VLDB 2026 Research / reviewers in the wild / expert
Yiyuan Yang
dblp:228/1875
· DBLP profile ↗
28ranked-venue papers
13as first author
27since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 10 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 8 first-author · 11 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Bi-level Personalization for Federated Foundation Models: A Task-vector Aggregation ApproachabstractFederated foundation models represent a new paradigm to jointly fine-tune pre-trained foundation models across clients. It is still a challenge to fine-tune foundation models for a small group of new users or specialized scenarios, which typically involve limited data compared to the large-scale data used in pre-training. In this context, the trade-off between personalization and federation becomes more sensitive. To tackle these, we proposed a bi-level personalization framework for federated fine-tuning on foundation models. Specifically, we conduct personalized fine-tuning on the client-level using its private data, and then conduct a personalized aggregation on the server-level using similar users measured by client-specific task vectors. Given the personalization information gained from client-level fine-tuning, the server-level personalized aggregation can gain group-wise personalization information while mitigating the disturbance of irrelevant or interest-conflict clients with non-IID data. The effectiveness of the proposed algorithm has been demonstrated by extensive experimental analysis in benchmark datasets. Yiyuan Yang, Guodong Long, Qinghua Lu 0001, Liming Zhu 0001, Jing Jiang 0002 |
AAAI | 1 |
| 2025 | AoI-MDP: An AoI Optimized Markov Decision Process Dedicated in the Underwater Task (Student Abstract)abstractOcean exploration places high demands on autonomous underwater vehicles, especially when there's observation delay. We propose age of information optimized Markov decision process (AoI-MDP) to enhance underwater tasks by modeling observation delay as signal delay and including it in the state space. AoI-MDP also introduces wait time in the action space and integrates AoI with reward functions, optimizing information freshness and decision-making using reinforcement learning. Simulations show AoI-MDP outperforms the standard MDP, demonstrating superior performance, feasibility, and generalization in underwater tasks. To accelerate relevant research, we have made the codes available as open-source at https://github.com/Xiboxtg/AoI-MDP. Yimian Ding, Jingzehua Xu, Yiyuan Yang, Guanwen Xie, Xinqi Wang, Shuai Zhang 0015 |
AAAI | 3 |
| 2025 | ERFSL: An Efficient Reward Function Searcher via Large Language Models for Custom-Environment Multi-Objective Reinforcement Learning (Student Abstract)abstractWe propose ERFSL, an efficient reward function searcher using large language models (LLMs) for custom-environment, multi-objective reinforcement learning (RL). ERFSL generates reward components based on explicit user requirements and rectifies them, and iteratively optimizes the weights of these components based on textual context. Applied to an underwater data collection RL task, ERFSL corrects reward codes with only one feedback iteration per requirement, and acquires diverse reward functions within the Pareto set. ERFSL also presents robust capability for deviated weights and small-size LLMs such as GPT-4o mini. The full-text prompts, examples of LLM-generated answers, and source code are available at https://360zmem.github.io/LLMRsearcher/ . Guanwen Xie, Jingzehua Xu, Yiyuan Yang, Yimian Ding, Shuai Zhang 0015 |
AAAI | 3 |
| 2025 | Time-MQA: Time Series Multi-Task Question Answering with Context EnhancementabstractYaxuan Kong, Yiyuan Yang, Yoontae Hwang, Wenjie Du, Stefan Zohren, Zhangyang Wang, Ming Jin, Qingsong Wen. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Yaxuan Kong, Yiyuan Yang, Yoontae Hwang, Stefan Zohren, Zhangyang Wang, Ming Jin 0005, Qingsong Wen |
ACL (1) | 2 |
| 2025 | Deep Learning for Multivariate Time Series Imputation: A SurveyabstractMissing values are ubiquitous in multivariate time series (MTS) data, posing significant challenges for accurate analysis and downstream applications. In recent years, deep learning-based methods have successfully handled missing data by leveraging complex temporal dependencies and learned data distributions. In this survey, we provide a comprehensive summary of deep learning approaches for multivariate time series imputation (MTSI) tasks. We propose a novel taxonomy that categorizes existing methods based on two key perspectives: imputation uncertainty and neural network architecture. Furthermore, we summarize existing MTSI toolkits with a particular emphasis on the PyPOTS Ecosystem, which provides an integrated and standardized foundation for MTSI research. Finally, we discuss key challenges and future research directions, which give insight for further MTSI research. This survey aims to serve as a valuable resource for researchers and practitioners in the field of time series analysis and missing data imputation tasks. A well-maintained MTSI paper and tool list is available at https://github.com/WenjieDu/Awesome_Imputation. Jun Wang 0121, Yiyuan Yang, Linglong Qian, Keli Zhang, Yuxuan Liang 0002, Qingsong Wen |
IJCAI | 3 |
| 2025 | Federated Low-Rank Adaptation for Foundation Models: A SurveyabstractEffectively leveraging private datasets remains a significant challenge in developing foundation models. Federated Learning (FL) has recently emerged as a collaborative framework that enables multiple users to fine-tune these models while mitigating data privacy risks. Meanwhile, Low-Rank Adaptation (LoRA) offers a resource-efficient alternative for fine-tuning foundation models by dramatically reducing the number of trainable parameters. This survey examines how LoRA has been integrated into federated fine-tuning for foundation models—an area we term FedLoRA—by focusing on three key challenges: distributed learning, heterogeneity, and efficiency. We further categorize existing work based on the specific methods used to address each challenge. Finally, we discuss open research questions and highlight promising directions for future investigation, outlining the next steps for advancing FedLoRA. Yiyuan Yang, Guodong Long, Qinghua Lu 0001, Liming Zhu 0001, Jing Jiang 0002, Chengqi Zhang |
IJCAI | 1 |
| 2025 | Efficient and Microphone-Fault-Tolerant 3D Sound Source Localization
Yiyuan Yang, Shitong Xu, Agathoniki Trigoni, Andrew Markham |
INTERSPEECH | 1 |
| 2025 | Target Speaker Extraction through Comparing Noisy Positive and Negative Audio EnrollmentsabstractTarget speaker extraction focuses on isolating a specific speaker's voice from an audio mixture containing multiple speakers. To provide information about the target speaker's identity, prior works have utilized clean audio samples as conditioning inputs. However, such clean audio examples are not always readily available. For instance, obtaining a clean recording of a stranger's voice at a cocktail party without leaving the noisy environment is generally infeasible. Limited prior research has explored extracting the target speaker's characteristics from noisy enrollments, which may contain overlapping speech from interfering speakers. In this work, we explore a novel enrollment strategy that encodes target speaker information from the noisy enrollment by comparing segments where the target speaker is talking (Positive Enrollments) with segments where the target speaker is silent (Negative Enrollments). Experiments show the effectiveness of our model architecture, which achieves over 2.1 dB higher SI-SNRi compared to prior works in extracting the monaural speech from the mixture of two speakers. Additionally, the proposed two-stage training strategy accelerates convergence, reducing the number of optimization steps required to reach 3 dB SNR by 60\%. Overall, our method achieves state-of-the-art performance in the monaural target speaker extraction conditioned on noisy enrollments. Our implementation is available at
https://github.com/xu-shitong/TSE-through-Positive-Negative-Enroll . Shitong Xu, Yiyuan Yang, Agathoniki Trigoni, Andrew Markham |
NeurIPS | 2 |
| 2025 | ScatterAD: Temporal-Topological Scattering Mechanism for Time Series Anomaly DetectionabstractOne main challenge in time series anomaly detection for industrial IoT lies in the complex spatio-temporal couplings within multivariate data. However, as traditional anomaly detection methods focus on modeling spatial or temporal dependencies independently, resulting in suboptimal representation learning and limited sensitivity to anomalous dispersion in high-dimensional spaces. In this work, we conduct an empirical analysis showing that both normal and anomalous samples tend to scatter in high-dimensional space, especially anomalous samples are markedly more dispersed. We formalize this dispersion phenomenon as scattering, quantified by the mean pairwise distance among sample representations, and leverage it as an inductive signal to enhance spatio-temporal anomaly detection. Technically, we propose ScatterAD to model representation scattering across temporal and topological dimensions. ScatterAD incorporates a topological encoder for capturing graph-structured scattering and a temporal encoder for constraining over-scattering through mean squared error minimization between neighboring time steps. We introduce a contrastive fusion mechanism to ensure the complementarity of the learned temporal and topological representations. Additionally, we theoretically show that maximizing the conditional mutual information between temporal and topological views improves cross-view consistency and enhances more discriminative representations. Extensive experiments on multiple public benchmarks show that ScatterAD achieves state-of-the-art performance on multivariate time series anomaly detection. Shaochen Fu, Li Huang 0006, Xiaohong Zhang 0002, Yiyuan Yang, Meng Yan 0001 |
NeurIPS | 6 |
| 2025 | PatchAD: A Lightweight Patch-Based MLP-Mixer for Time Series Anomaly DetectionabstractTime series anomaly detection is a pivotal task in data analysis, yet it poses the challenge of discerning normal and abnormal patterns in label-deficient scenarios. While prior studies have largely employed reconstruction-based approaches, which limit the models’ representational capacities. Moreover, existing deep learning-based methods are not sufficiently lightweight. Addressing these issues, we present PatchAD, our novel, highly efficient multiscale patch-based MLP-Mixer architecture that utilizes contrastive learning for representation extraction and anomaly detection. With its four distinct MLP Mixers and innovative dual project constraint module, PatchAD mitigates potential model degradation and offers a lightweight solution, requiring only0.403 Mparameters. Its efficacy is demonstrated by state-of-the-art results across8datasets sourced from different application scenarios, outperforming over30comparative algorithms. PatchAD significantly improves the classical F1 score by6.84%, the Aff-F1 score by4.27%, and the V-ROC by2.49%. Simultaneously, an in-depth analysis of the mechanisms underlying PatchAD has been conducted from both theoretical and experimental perspectives, validating the design motivations of the model. Zhiwen Yu 0002, Yiyuan Yang, Weizheng Wang 0001, Kaixiang Yang 0001, C. L. Philip Chen |
IEEE Trans. Big Data | 3 |
| 2025 | DACAD: Domain Adaptation Contrastive Learning for Anomaly Detection in Multivariate Time SeriesabstractIn time series anomaly detection (TSAD), the scarcity of labeled data poses a challenge to the development of accurate models. Unsupervised domain adaptation (UDA) offers a solution by leveraging labeled data from a related domain to detect anomalies in an unlabeled target domain. However, existing UDA methods assume consistent anomalous classes across domains. To address this limitation, we propose a novel Domain Adaptation Contrastive learning model for Anomaly Detection in multivariate time series (DACAD), combining UDA with contrastive learning. DACAD utilizes an anomaly injection mechanism that enhances generalization across unseen anomalous classes, improving adaptability and robustness. Additionally, our model employs supervised contrastive loss for the source domain and self-supervised contrastive triplet loss for the target domain, ensuring comprehensive feature representation learning and domain-invariant feature extraction. Finally, an effective Center-based Entropy Classifier (CEC) accurately learns normal boundaries in the source domain. Extensive evaluations on multiple real-world datasets and a synthetic dataset highlight DACAD's superior performance in transferring knowledge across domains and mitigating the challenge of limited labeled data in TSAD. Zahra Zamanzadeh Darban, Yiyuan Yang, Geoffrey I. Webb, Charu C. Aggarwal, Qingsong Wen, Shirui Pan, Mahsa Salehi |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2025 | SimAD: A Simple Dissimilarity-Based Approach for Time-Series Anomaly DetectionabstractDespite the prevalence of reconstruction-based deep learning methods, time-series anomaly detection (TSAD) remains a tremendous challenge. Existing approaches often struggle with limited temporal contexts, insufficient representation of normal patterns, and flawed evaluation metrics, all of which hinder their effectiveness in detecting anomalous behavior. To address these issues, we introduce a simple dissimilarity-based approach for time-series anomaly detection (SimAD). Specifically, SimAD first incorporates a patching-based feature extractor capable of processing extended temporal windows and employs the EmbedPatch encoder to fully integrate normal behavioral patterns. Second, we design an innovative ContrastFusion module in SimAD, which strengthens the robustness of anomaly detection by highlighting the distributional differences between normal and abnormal data. Third, we introduce two robust enhanced evaluation metrics, unbiased affiliation (UAff) and normalized affiliation (NAff), designed to overcome the limitations of existing metrics by providing better distinctiveness and semantic clarity. The reliability of these two metrics has been demonstrated by both theoretical and experimental analyses. Experiments conducted on seven diverse time-series datasets clearly demonstrate SimAD's superior performance compared with state-of-the-art (SOTA) methods, achieving relative improvements of 19.85% on ${F}1$ , 4.44% on Aff-F1, 77.79% on NAff-F1, and 9.69% on AUC on six multivariate datasets. Code and pretrained models are available at https://github.com/EmorZz1G/SimAD. Zhiwen Yu 0002, Xing Xi, Wenming Cao 0002, Yiyuan Yang, Kaixiang Yang 0001, Jane You |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | AI-Based Energy Transportation Safety: Pipeline Radial Threat Estimation Using Intelligent Sensing SystemabstractThe application of artificial intelligence technology has greatly enhanced and fortified the safety of energy pipelines, particularly in safeguarding against external threats. The predominant methods involve the integration of intelligent sensors to detect external vibration, enabling the identification of event types and locations, thereby replacing manual detection methods. However, practical implementation has exposed a limitation in current methods - their constrained ability to accurately discern the spatial dimensions of external signals, which complicates the authentication of threat events. Our research endeavors to overcome the above issues by harnessing deep learning techniques to achieve a more fine-grained recognition and localization process. This refinement is crucial in effectively identifying genuine threats to pipelines, thus enhancing the safety of energy transportation. This paper proposes a radial threat estimation method for energy pipelines based on distributed optical fiber sensing technology. Specifically, we introduce a continuous multi-view and multi-domain feature fusion methodology to extract comprehensive signal features and construct a threat estimation and recognition network. The utilization of collected acoustic signal data is optimized, and the underlying principle is elucidated. Moreover, we incorporate the concept of transfer learning through a pre-trained model, enhancing both recognition accuracy and training efficiency. Empirical evidence gathered from real-world scenarios underscores the efficacy of our method, notably in its substantial reduction of false alarms and remarkable gains in recognition accuracy. More generally, our method exhibits versatility and can be extrapolated to a broader spectrum of recognition tasks and scenarios. Chengyuan Zhu, Yiyuan Yang, Kaixiang Yang 0001, Qinmin Yang, C. L. Philip Chen |
AAAI | 2 |
| 2024 | Advancing Multivariate Time Series Anomaly Detection: A Comprehensive Benchmark with Real-World Data from Alibaba CloudabstractTime series anomaly detection is of significant importance in many real-world applications, including finance, healthcare, network security, industrial equipment, complex computing systems, and space probes. Most of these applications involve multi-sensor systems, thus how to perform multivariate time series anomaly detection (MTSAD) has garnered widespread attention. This broad attention has fueled extensive research endeavors aimed to innovate and develop methods and techniques to improve the efficiency and precision of anomaly detection on multivariate time series data, including both classic machine learning methods and deep learning methods. However, evaluating the performance of these methods remains challenging due to the limited availability of public benchmark datasets for MTSAD, which are often criticized for various reasons. Additionally, there is no consensus on the best metrics for time series anomaly detection, further complicating MTSAD research. In this paper, we advance the benchmarking of time series anomaly detection by addressing datasets, evaluation metrics, and algorithm comparison. To the best of our knowledge, we have generated the largest real-world datasets for MTSAD using the Hologres AIOps system in the Alibaba Cloud platform. We review and compare popular evaluation metrics including recently proposed ones. To evaluate classic machine learning and recent deep learning methods fairly, we have conducted extensive comparisons of these methods on various datasets. We believe that our benchmarks and datasets will promote reproducible results and accelerate the progress of MTSAD research. Chaoli Zhang 0001, Lanshu Peng, Qingsong Wen, Yiyuan Yang, Chong-Jiong Fan, Minqi Jiang, Lunting Fan, Liang Sun 0001 |
CIKM | 5 |
| 2024 | Pre-training Cross-Modal Retrieval by Expansive Lexicon-Patch AlignmentabstractRecent large-scale vision-language pre-training depends on image-text global alignment by contrastive learning and is further boosted by fine-grained alignment in a weakly contrastive manner for cross-modal retrieval. Nonetheless, besides semantic matching learned by contrastive learning, cross-modal retrieval also largely relies on object matching between modalities. This necessitates fine-grained categorical discriminative learning, which however suffers from scarce data in full-supervised scenarios and information asymmetry in weakly-supervised scenarios when applied to cross-modal retrieval. To address these issues, we propose expansive lexicon-patch alignment (ELA) to align image patches with a vocabulary rather than only the words explicitly in the text for annotation-free alignment and information augmentation, thus enabling more effective fine-grained categorical discriminative learning for cross-modal retrieval. Experimental results show that ELA could effectively learn representative fine-grained information and outperform state-of-the-art methods on cross-modal retrieval. Yiyuan Yang, Guodong Long, Michael Blumenstein, Xiubo Geng, Chongyang Tao, Tao Shen 0001, Daxin Jiang |
LREC/COLING | 1 |
| 2024 | Analyzing Women's Contributions to Open-Source Software Projects based on Large Language ModelsabstractOpen-source software (OSS) enables users to access, modify, distribute software based on open-source licenses, serving as vital digital infrastructure. Notably, GitHub stands out as a prominent OSS community, with 94 million developers engaged in projects by 2022. However, accurately assessing women’s contributions in OSS encounters challenges due to limited gender data. To address this, we propose an innovative method that employs the Large-Language-Model (LLM), ChatLM2. This LLM-based approach allows cross-lingual analysis of women’s involvement and quantitatively assesses their impact on OSS projects. The study aims to uncover gender disparities and encourage greater participation of female developers in the open-source realm. The article is structured with sections on research methods, design, LLM-based gender detection, women’s participation, impact assessment, implications, and future research. Yuqian Zhuang, Mingya Zhang, Yiyuan Yang |
CSCWD | 3 |
| 2024 | SSL-Net: A Synergistic Spectral and Learning-Based Network for Efficient Bird Sound ClassificationabstractEfficient and accurate bird sound classification is of important for ecology, habitat protection and scientific research, as it plays a central role in monitoring the distribution and abundance of species. However, prevailing methods typically demand extensively labeled audio datasets and have highly customized frameworks, imposing substantial computational and annotation loads. In this study, we present an efficient and general framework called SSL-Net, which combines spectral and learned features to identify different bird sounds. Encouraging empirical results gleaned from a standard field-collected bird audio dataset validate the efficacy of our method in extracting features efficiently and achieving heightened performance in bird sound classification, even when working with limited sample sizes. Furthermore, we present three feature fusion strategies, aiding engineers and researchers in their selection through quantitative analysis. Yiyuan Yang, Kaichen Zhou, Agathoniki Trigoni, Andrew Markham |
ICASSP | 1 |
| 2024 | Pre-training Feature Guided Diffusion Model for Speech Enhancement
Yiyuan Yang, Agathoniki Trigoni, Andrew Markham |
INTERSPEECH | 1 |
| 2024 | Dual-Personalizing Adapter for Federated Foundation ModelsabstractRecently, foundation models, particularly large language models (LLMs), have demonstrated an impressive ability to adapt to various tasks by fine-tuning diverse instruction data. Notably, federated foundation models (FedFM) emerge as a privacy preservation method to fine-tune models collaboratively under federated learning (FL) settings by leveraging many distributed datasets with non-IID data. To alleviate communication and computation overhead, parameter-efficient methods are introduced for efficiency, and some research adapted personalization methods to FedFM for better user preferences alignment. However, a critical gap in existing research is the neglect of test-time distribution shifts in real-world applications, and conventional methods for test-time distribution shifts in personalized FL are less effective for FedFM due to their failure to adapt to complex distribution shift scenarios and the requirement to train all parameters. To bridge this gap, we refine the setting in FedFM, termed test-time personalization, which aims to learn personalized federated foundation models on clients while effectively handling test-time distribution shifts simultaneously. To address challenges in this setting, we explore a simple yet effective solution, a Federated Dual-Personalizing Adapter (FedDPA) architecture. By co-working with a foundation model, a global adapter and a local adapter jointly tackle the test-time distribution shifts and client-specific personalization. Additionally, we introduce an instance-wise dynamic weighting mechanism that dynamically integrates the global and local adapters for each test instance during inference, facilitating effective test-time personalization. The effectiveness of the proposed method has been evaluated on benchmark datasets across different NLP tasks. Yiyuan Yang, Guodong Long, Tao Shen 0001, Jing Jiang 0002, Michael Blumenstein |
NeurIPS | 1 |
| 2024 | Localizing and tracking of in-pipe inspection robots based on distributed optical fiber sensing
Chengyuan Zhu, Yanyun Pu, Yiyuan Yang, Zhuoling Lyu, Chao Li 0062, Qinmin Yang |
Adv. Eng. Informatics | 3 |
| 2023 | SGDP: A Stream-Graph Neural Network Based Data PrefetcherabstractData prefetching is important for storage system optimization and access performance improvement. Traditional prefetchers work well for mining access patterns of sequential logical block address (LBA) but cannot handle complex non-sequential patterns that commonly exist in real-world applications. The state-of-the-art (SOTA) learning-based prefetchers cover more LBA accesses. However, they do not adequately consider the spatial interdependencies between LBA deltas, which leads to limited performance and robustness. This paper proposes a novel Stream-Graph neural network-based Data Prefetcher (SGDP). Specifically, SGDP models LBA delta streams using a weighted directed graph structure to represent interactive relations among LBA deltas and further extracts hybrid features by graph neural networks for data prefetching. We conduct extensive experiments on eight real-world datasets. Empirical results verify that SGDP outperforms the SOTA methods in terms of the hit ratio by 6.21%, the effective prefetching ratio by 7.00%, and speeds up inference time by 3.13× on average. Besides, we generalize SGDP to different variants by different stream constructions, further expanding its application scenarios and demonstrating its robustness. SGDP offers a novel data prefetching solution and has been verified in commercial hybrid storage systems in the experimental phase. Our codes and appendix are available at https://github.com/yyysjz1997/SGDP/. Yiyuan Yang, Rongshang Li, Qiquan Shi, Xijun Li, Xing Li 0023, Mingxuan Yuan |
IJCNN | 1 |
| 2023 | DCdetector: Dual Attention Contrastive Representation Learning for Time Series Anomaly DetectionabstractTime series anomaly detection is critical for a wide range of applications. It aims to identify deviant samples from the normal sample distribution in time series. The most fundamental challenge for this task is to learn a representation map that enables effective discrimination of anomalies. Reconstruction-based methods still dominate, but the representation learning with anomalies might hurt the performance with its large abnormal loss. On the other hand, contrastive learning aims to find a representation that can clearly distinguish any instance from the others, which can bring a more natural and promising representation for time series anomaly detection. In this paper, we propose DCdetector, a multi-scale dual attention contrastive representation learning model. DCdetector utilizes a novel dual attention asymmetric design to create the permutated environment and pure contrastive loss to guide the learning process, thus learning a permutation invariant representation with superior discrimination abilities. Extensive experiments show that DCdetector achieves state-of-the-art results on multiple time series anomaly detection benchmark datasets. Code is publicly available at https://github.com/DAMO-DI-ML/KDD2023-DCdetector. Yiyuan Yang, Chaoli Zhang 0001, Tian Zhou 0004, Qingsong Wen, Liang Sun 0001 |
KDD | 1 |
| 2023 | DynPoint: Dynamic Neural Point For View SynthesisabstractThe introduction of neural radiance fields has greatly improved the effectiveness of view synthesis for monocular videos. However, existing algorithms face difficulties when dealing with uncontrolled or lengthy scenarios, and require extensive training time specific to each new scenario.
To tackle these limitations, we propose DynPoint, an algorithm designed to facilitate the rapid synthesis of novel views for unconstrained monocular videos.
Rather than encoding the entirety of the scenario information into a latent representation, DynPoint concentrates on predicting the explicit 3D correspondence between neighboring frames to realize information aggregation.
Specifically, this correspondence prediction is achieved through the estimation of consistent depth and scene flow information across frames.
Subsequently, the acquired correspondence is utilized to aggregate information from multiple reference frames to a target frame, by constructing hierarchical neural point clouds.
The resulting framework enables swift and accurate view synthesis for desired views of target frames.
The experimental results obtained demonstrate the considerable acceleration of training time achieved - typically an order of magnitude - by our proposed method while yielding comparable outcomes compared to prior approaches. Furthermore, our method exhibits strong robustness in handling long-duration videos without learning a canonical representation of video content. Kaichen Zhou, Jia-Xing Zhong, Sang-Yun Shin, Kai Lu 0003, Yiyuan Yang, Andrew Markham, Agathoniki Trigoni |
NeurIPS | 5 |
| 2021 | Early Safety Warnings for Long-Distance Pipelines: A Distributed Optical Fiber Sensor Machine Learning ApproachabstractAutomated pipeline safety early warning (PSEW) systems are designed to automatically identify and locate third-party damage events on oil and gas pipelines. They are intended to replace traditional, inefficient manual inspection methods. However, current PSEW methods cannot achieve universality for various complex environments because they are sensitive to the spatiotemporal stability of the signal obtained by its distributed sensors at various locations and times. Our research aimed to improve the accuracy of long-distance oil–gas PSEW systems through machine learning. In this paper, we propose a novel real-time action recognition method for long-distance PSEW systems based on a coherent Rayleigh scattering distributed optical fiber sensor. More specifically, we put forward two complementary feature calculation methods to describe signals and build a new action recognition deep learning network based on those features. Encouraging empirical results on the data collected at a real location confirm that the features can effectively describe signals in an environment with strong noise and weak signals, and the entire approach can identify and locate third-party damage events quickly under various hardware conditions with accuracies of 99.26% (500 Hz) and 97.20% (100 Hz). More generically, our method can be applied to other fields as well. Yiyuan Yang, Taojia Zhang |
AAAI | 1 |
| 2021 | Block Access Pattern Discovery via Compressed Full Tensor TransformerabstractThe discovery and prediction of block access patterns in hybrid storage systems is of crucial importance for effective tier management. Existing methods are usually based on heuristics and unable to handle complex patterns. This work newly introduces transformer to block access pattern prediction. We remark that block accesses in the tier management systems are aggregated temporally and spatially as multivariate time series of block access frequency, so the runtime requirements are relaxed, making complex models applicable for the deployment. Moreover, enormous and rarely accessed blocks in storage systems and the structure of traditional transformer models would result in millions of redundant parameters and make them impractical to be deployed. We incorporate Tensor-Train Decomposition (TTD) with transformer and propose the Compressed Full Tenor Transformer (CFTT), in which all linear layers in the vanilla transformer are replaced with tensor-train layers. Weights of input and output layers are shared to further reduce parameters and reuse knowledge implicitly. CFTT can significantly reduce the model size and computation cost, which is critical to save storage space and inference time. Extensive experiments are conducted on synthetic and real-world datasets. The results demonstrate that transformers achieve state-of-the-art performance stably in terms of top-k hit rates. Moreover, the proposed CFTT compresses transformers 16× to 461× and speeds up inference 5× without sacrificing performance on the whole, which facilitates its applications in tier management in hybrid storage systems. Xing Li 0023, Qiquan Shi, Lei Chen 0031, Yiyuan Yang, Mingxuan Yuan |
CIKM | 6 |
| 2021 | Do Time Constraints Re-Prioritize Attention to Shapes During Visual Photo Inspection?
Yiyuan Yang, Kenneth Li 0001, Fernanda Monteiro Eliott, Maithilee Kunda |
CogSci | 1 |
| 2021 | Pipeline Safety Early Warning Method for Distributed Signal using Bilinear CNN and LightGBMabstractOil and gas pipelines are known as the backbone of global energy, and securing their safety is crucial for energy supply. In this study, we utilized a novel machine learning method based on the spatiotemporal features of distributed optical fiber sensor signals to monitor the safety of oil and gas pipelines in real time. Encouraging empirical results on a large amount of data collected from real sites confirmed that our model could accurately locate and identify the damage events of a pipeline in real time under strong noise and various hardware conditions, and could effectively handle the signal drift problem. Furthermore, as a generalized tool, the proposed solution could be applied to other industrial inspection fields. Our codes and video demos are available at https://github.com/yyysjz1997/B-CNN_LGBM-PSEW. Yiyuan Yang |
ICASSP | 1 |
| 2019 | High capacity and multilevel information hiding algorithm based on pu partition modes for HEVC videos
Yiyuan Yang, Zhaohong Li, Wenchao Xie |
Multim. Tools Appl. | 1 |