EDBT 2026 Demo / reviewers in the wild / expert
Baoquan Zhang
dblp:160/1111
· DBLP profile ↗
57ranked-venue papers
17as first author
48since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 31 · 9 first-author · 31 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 6 first-author · 16 since 2021Systems, architecture and hardware · 12 · 4 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 3 first-author · 9 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Satellite-Text-Prompted Large Language Model for Photovoltaic Power ForecastingabstractPhotovoltaic (PV) power forecasting is critical for the operation of solar power plants and the coordination of energy within power grids. This work aims to predict future PV power time series by leveraging multimodal data. While recent studies have incorporated numerical modalities such as satellite image sequences and numerical weather prediction (NWP) time series, they often overlook textual modalities—such as the spatio-temporal context of PV plants—and the potential of pretrained large language models (LLMs). In this paper, we build upon existing numerical inputs and further explore the use of spatio-temporal text prompts, generated based on plant coordinates and forecast start time, to enhance the forecasting process. We propose PV-LLM, a satellite-text-prompted framework that integrates a pretrained LLM to improve PV power forecasting. The framework consists of three key components: Text Prompt Construction, Modality-Specific Encoding, and Adaptive Prompt Tuning. First, the Text Prompt Construction module generates spatio-temporal prompts that offer high-level semantic guidance. Next, the Modality-Specific Encoding module encodes each modality according to its unique characteristics, capturing modality-specific patterns while managing varying context lengths. Finally, the Adaptive Prompt Tuning module fine-tunes the LLM to integrate multimodal embeddings, while an adaptive gating mechanism retains its pretrained knowledge. We validate the effectiveness of the proposed framework on a real-world dataset containing multiple PV plants. Experimental results demonstrate that our approach outperforms existing state-of-the-art methods. Jianghong Ma, Baoquan Zhang, Kenghong Lin, Chuyao Luo, Xutao Li 0001, Yunming Ye |
AAAI | 3 |
| 2026 | Improved Masked Image Generation with Knowledge-Augmented Token RepresentationsabstractMasked image generation (MIG) has demonstrated remarkable efficiency and high-fidelity images by enabling parallel token prediction. Existing methods typically rely solely on the model itself to learn semantic dependencies among visual token sequences. However, directly learning such semantic dependencies from data is challenging because the individual tokens lack clear semantic meanings, and these sequences are usually long. To address this limitation, we propose a novel Knowledge-Augmented Masked Image Generation framework, named KA-MIG, which introduces explicit knowledge of token-level semantic dependencies (i.e., extracted from the training data) as priors to learn richer representations for improving performance. In particular, we explore and identify three types of advantageous token knowledge graphs, including two positive and one negative graphs (i.e., the co-occurrence graph, the semantic similarity graph, and the position-token incompatibility graph). Based on three prior knowledge graphs, we design a graph-aware encoder to learn token and position-aware representations. After that, a lightweight fusion mechanism is introduced to integrate these enriched representations into the existing MIG methods. Resorting to such prior knowledge, our method effectively enhances the model's ability to capture semantic dependencies, leading to improved generation quality. Experimental results demonstrate that our method improves upon existing MIG for class-conditional image generation on ImageNet. Guotao Liang, Baoquan Zhang, Zihao Han, Yunming Ye |
AAAI | 2 |
| 2026 | LMcast: A pretrained language model guided long-term memory transformer for precipitation nowcasting
Feifan Gao, Chuyao Luo, Guangbo Deng, Xutao Li 0003, Baoquan Zhang, Demin Yu, Yunming Ye |
Neural Networks | 5 |
| 2026 | Codebook Transfer With Vision-to-Language Translation for Vector Quantization
Baoquan Zhang, Guotao Liang, Tianran Chen, Yunming Ye, Xiaochen Qi |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | Advection-diffusion spatiotemporal recurrent network for regional wind speed prediction
Shidong Chen, Baoquan Zhang, Xutao Li 0003, Yunming Ye, Kenghong Lin, Rui Ye 0002 |
Pattern Recognit. | 2 |
| 2025 | AsyncDSB: Schedule-Asynchronous Diffusion Schrödinger Bridge for Image InpaintingabstractImage inpainting is an important image generation task, which aims to restore corrupted image from partial visible area. Recently, diffusion Schrödinger bridge methods effectively tackle this task by modeling the translation between corrupted and target images as a diffusion Schrödinger bridge process along a noising schedule path. Although these methods have shown superior performance, in this paper, we find that 1) existing methods suffer from a schedule-restoration mismatching issue, i.e., the theoretical schedule and practical restoration processes usually exist a large discrepancy, which theoretically results in the schedule not fully leveraged for restoring images; and 2) the key reason causing such issue is that the restoration process of all pixels are actually asynchronous but existing methods set a synchronous noise schedule to them, i.e., all pixels shares the same noise schedule. To this end, we propose a schedule-Asynchronous Diffusion Schrödinger Bridge (AsyncDSB) for image inpainting. Our insight is preferentially scheduling pixels with high frequency (i.e., large gradients) and then low frequency (i.e., small gradients). Based on this insight, given a corrupted image, we first train a network to predict its gradient map in corrupted area. Then, we regard the predicted image gradient as prior and design a simple yet effective pixel-asynchronous noise schedule strategy to enhance the diffusion Schrödinger bridge. Thanks to the asynchronous schedule at pixels, the temporal interdependence of restoration process between pixels can be fully characterized for high-quality image inpainting. Experiments on real-world datasets show that our AsyncDSB achieves superior performance, especially on FID with around 3% ∼ 14% improvement over state-of-the-art baseline methods. Zihao Han, Baoquan Zhang, Lisai Zhang, Shanshan Feng 0001, Kenghong Lin, Guotao Liang, Yunming Ye, Joeq, Kola Ye |
AAAI | 2 |
| 2025 | Sensitivity-Aware Efficient Fine-Tuning via Compact Dynamic-Rank AdaptationabstractParameter-Efficient Fine-Tuning (PEFT) is a fundamental research problem in computer vision, which aims to tune a few parameters for efficient storage and adaptation of pre-trained vision models. Recently, sensitivity-aware parameter efficient fine-tuning method (SPT) addresses this problem by identifying sensitive parameters and then leveraging its sparse characteristic to combine unstructured and structured tuning for PEFT. However, existing methods only focus on the sparse characteristic of sensitive parameters but overlook its distribution characteristic, which results in additional storage burden and limited performance improvement. In this paper, we find that the distribution of sensitive parameters is not chaotic, but concentrates on a small number of rows or columns in each parameter matrix. Inspired by this fact, we propose a Compact Dynamic-Rank Adaptation-based tuning method for Sensitivity-aware Parameter efficient fine-Tuning, called CDRA-SPT. Specifically, we first identify the sensitive parameters that require tuning for each down-stream task. Then, we reorganize the sensitive parameters by following its row and column into a compact sub-parameter matrix. Finally, a dynamic-rank adaptation is designed and applied at sub-parameter matrix level for PEFT. Its advantage is that the dynamic-rank characteristic of sub-parameter matrix can be fully exploited for PEFT. Extensive experiments show that our method achieves superior performance over previous state-of-the-art methods. Tianran Chen, Jiarui Chen, Baoquan Zhang, Zhehao Yu, Shidong Chen, Rui Ye 0002, Xutao Li 0003, Yunming Ye |
CVPR | 3 |
| 2025 | Towards Improved Text-Aligned Codebook Learning: Multi-Hierarchical Codebook-Text Alignment with Long TextabstractImage quantization is a crucial technique in image generation, aimed at learning a codebook that encodes an image into a discrete token sequence. Recent advancements have seen researchers exploring learning multi-modal codebook (i.e., text-aligned codebook) by utilizing image caption semantics, aiming to enhance codebook performance in cross-modal tasks. However, existing image-text paired datasets exhibit a notable flaw in that the text descriptions tend to be overly concise, failing to adequately describe the images and provide sufficient semantic knowledge, resulting in limited alignment of text and codebook at a fine-grained level. In this paper, we propose a novel Text-Augmented Codebook Learning framework, named TA-VQ, which generates longer text for each image using the visual-language model for improved text-aligned codebook learning. However, the long text presents two key challenges: how to encode text and how to align codebook and text. To tackle two challenges, we propose to split the long text into multiple granularities for encoding, i.e., word, phrase, and sentence, so that the long text can be fully encoded without losing any key semantic knowledge. Following this, a hierarchical encoder and novel sampling-based alignment strategy are designed to achieve fine-grained codebook-text alignment. Additionally, our method can be seamlessly integrated into existing VQ models. Extensive experiments in reconstruction and various downstream tasks demonstrate its effectiveness compared to previous state-of-the-art approaches. Guotao Liang, Baoquan Zhang, Junteng Zhao, Yunming Ye, Kola Ye |
CVPR | 2 |
| 2025 | AlphaPre: Amplitude-Phase Disentanglement Model for Precipitation NowcastingabstractPrecipitation nowcasting involves using current radar observation sequences to predict future radar sequences and determine future precipitation distribution, which is crucial for disaster warning, traffic planning, and agricultural production. Despite numerous advancements, challenges persist in accurately predicting both the location and intensity of precipitation, as these factors are often interdependent, with complex atmospheric dynamics and moisture distribution causing position and intensity changes to be intricately coupled. Inspired by the fact that in the frequency domain, phase variations are shown to correspond to changes in the position of precipitation, while amplitude variations are linked to intensity changes, we propose an amplitude-phase disentanglement model called AlphaPre, which separately learn the position and intensity changes of precipitation. AlphaPre comprises three key components: a phase network, an amplitude network, and an AlphaMixer. The phase network captures positional changes by learning phase variations, and the amplitude network models intensity changes by alternating between the frequency and spatial domains. The AlphaMixer then integrates these components to produce a refined precipitation forecast. Extensive experiments on four datasets demonstrate the effectiveness and superiority of our method over state-of-the-art approaches. Our code is publicly available at https://github.com/linkenghong/AlphaPre. Kenghong Lin, Baoquan Zhang, Demin Yu, Wenzhi Feng, Shidong Chen, Feifan Gao, Xutao Li 0003, Yunming Ye |
CVPR | 2 |
| 2025 | DEEPSERVE: Serverless Large Language Model Serving at Scale
Zhixia Liu, Yuetao Chen, Baoquan Zhang, Shining Wan, Gengyuan Dan, Zhiyu Dong, Zhihao Ren, Changhong Liu, Tao Xie 0001, Dayun Lin, Xusheng Chen, Yizhou Shan |
USENIX ATC | 9 |
| 2025 | ZooKT: Task-adaptive knowledge transfer of Model Zoo for few-shot learning
Baoquan Zhang, Bingqi Shan, Aoxue Li, Chuyao Luo, Yunming Ye, Zhenguo Li |
Pattern Recognit. | 1 |
| 2025 | Knowledge-Augmented Interpretable Network for Zero-Shot Stance Detection on Social MediaabstractStance detection on social media has become increasingly important for understanding public opinions on controversial issues. Existing methods often require large amounts of labeled data to learn target-independent transferable knowledge, which is infeasible under zero-shot settings where the target is unseen. Furthermore, most current stance detection models, primarily based on end-to-end deep learning architectures, lack transparency and may produce counter-intuitive and uninterpretable predictions. In this article, we propose a novel knowledge-augmented interpretable network (KAI) to enable zero-shot stance detection (ZSSD). First, we introduce an unsupervised approach based on large language models (LLMKE) to elicit analysis perspectives, which is target-independent knowledge shared across different targets. This transferable knowledge bridges connections between seen and unseen targets. Second, we develop a bidirectional knowledge-guided neural production system (Bi-KGNPS) that effectively integrates such transferable knowledge through an iterative knowledge-variable binding process to guide stance predictions. Extensive experiments on benchmark datasets demonstrate KAI achieves new state-of-the-art performance on ZSSD. Moreover, our approach also delivers strong results on conventional in-target and cross-target stance detection. With the dual benefits of knowledge-augmented accuracy and model interpretability, this work represents an important advance toward practical stance detection systems that can generalize to emerging topics of interest. The proposed KAI framework provides an interpretable approach to effectively transfer knowledge across domains for zero-shot learning. Bowen Zhang 0005, Daijun Ding, Zhichao Huang 0001, Ang Li 0047, Baoquan Zhang, Hu Huang 0009 |
IEEE Trans. Comput. Soc. Syst. | 6 |
| 2025 | RePA: Rebalance Parameter Adaptation for Precipitation Nowcasting
Youran Wang, Baoquan Zhang, Yunming Ye, Xutao Li 0003 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | HPCR: Holistic Proxy-Based Contrastive Replay for Online Continual LearningabstractOnline continual learning (OCL), aimed at developing a neural network that continuously learns new data from a single pass over an online data stream, generally suffers from catastrophic forgetting (CF). Existing replay-based methods alleviate forgetting by replaying partial old data in a proxy-based or contrastive-based replay manner, each with its own shortcomings. Our previous work proposes a novel replay-based method called proxy-based contrastive replay (PCR), which handles the shortcomings by achieving complementary advantages of both replay manners. In this work, we further conduct gradient and limitation analysis of PCR. The analysis results show that PCR still can be further improved in feature extraction, generalization, and anti-forgetting capabilities of the model. Hence, we developed a more advanced method named holistic PCR (HPCR). HPCR consists of three components, each tackling one of the limitations of PCR. The contrastive component conditionally incorporates anchor-to-sample pairs to PCR, improving the feature extraction ability. The second is a temperature component that decouples the temperature coefficient into two parts based on their gradient impacts and sets different values for them to enhance the generalization ability. The third is a distillation component that constrains the learning process with additional loss terms to improve the anti-forgetting ability. Experiments on four datasets consistently demonstrate the superiority of HPCR over various state-of-the-art methods. Huiwei Lin, Shanshan Feng 0001, Baoquan Zhang, Xutao Li 0003, Yunming Ye |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | MetaDiff: Meta-Learning with Conditional Diffusion for Few-Shot LearningabstractEquipping a deep model the ability of few-shot learning (FSL) is a core challenge for artificial intelligence. Gradient-based meta-learning effectively addresses the challenge by learning how to learn novel tasks. Its key idea is learning a deep model in a bi-level optimization manner, where the outer-loop process learns a shared gradient descent algorithm (called meta-optimizer), while the inner-loop process leverages it to optimize a task-specific base learner with few examples. Although these methods have shown superior performance on FSL, the outer-loop process requires calculating second-order derivatives along the inner-loop path, which imposes considerable memory burdens and the risk of vanishing gradients. This degrades meta-learning performance. Inspired by recent diffusion models, we find that the inner-loop gradient descent process can be viewed as a reverse process (i.e., denoising) of diffusion where the target of denoising is the weight of base learner but origin data. Based on this fact, we propose to model the gradient descent algorithm as a diffusion model and then present a novel conditional diffusion-based meta-learning, called MetaDiff, that effectively models the optimization process of base learner weights from Gaussian initialization to target weights in a denoising manner. Thanks to the training efficiency of diffusion models, our MetaDiff does not need to differentiate through the inner-loop path such that the memory burdens and the risk of vanishing gradients can be effectively alleviated for improving FSL. Experimental results show that our MetaDiff outperforms state-of-the-art gradient-based meta-learning family on FSL tasks. Baoquan Zhang, Chuyao Luo, Demin Yu, Xutao Li 0003, Huiwei Lin, Yunming Ye, Bowen Zhang 0005 |
AAAI | 1 |
| 2024 | A Challenge Dataset and Effective Models for Conversational Stance DetectionabstractPrevious stance detection studies typically concentrate on evaluating stances within individual instances, thereby exhibiting limitations in effectively modeling multi-party discussions concerning the same specific topic, as naturally transpire in authentic social media interactions. This constraint arises primarily due to the scarcity of datasets that authentically replicate real social media contexts, hindering the research progress of conversational stance detection. In this paper, we introduce a new multi-turn conversation stance detection dataset (called MT-CSD), which encompasses multiple targets for conversational stance detection. To derive stances from this challenging dataset, we propose a global-local attention network (GLAN) to address both long and short-range dependencies inherent in conversational data. Notably, even state-of-the-art stance detection methods, exemplified by GLAN, exhibit an accuracy of only 50.47%, highlighting the persistent challenges in conversational stance detection. Furthermore, our MT-CSD dataset serves as a valuable resource to catalyze advancements in cross-domain stance detection, where a classifier is adapted from a different yet related target. We believe that MT-CSD will contribute to advancing real-world applications of stance detection research. Our source code, data, and models are available at https://github.com/nfq729/MT-CSD. Fuqiang Niu, Min Yang 0007, Ang Li 0047, Baoquan Zhang, Xiaojiang Peng, Bowen Zhang 0005 |
LREC/COLING | 4 |
| 2024 | DiffCast: A Unified Framework via Residual Diffusion for Precipitation NowcastingabstractPrecipitation nowcasting is an important spatiotemporal prediction task to predict the radar echoes sequences based on current observations, which can serve both meteorological science and smart city applications. Due to the chaotic evolution nature of the precipitation systems, it is a very challenging problem. Previous studies address the problem either from the perspectives of deterministic modeling or probabilistic modeling. However, their predictions suffer from the blurry, high-value echoes fading away and position inaccurate issues. The root reason of these issues is that the chaotic evolutionary precipitation systems are not appropriately modeled. Inspired by the nature of the systems, we propose to decompose and model them from the perspective of global deterministic motion and local stochastic variations with residual mechanism. A unified and flexible framework that can equip any type of spatio-temporal models is proposed based on residual diffusion, which effectively tackles the shortcomings of previous methods. Extensive experimental results on four publicly available radar datasets demonstrate the effectiveness and superiority of the proposed framework, compared to state-of-the-art techniques. Our code is publicly available at https://github.com/DeminYu98/DiffCast. Demin Yu, Xutao Li 0003, Yunming Ye, Baoquan Zhang, Chuyao Luo, Kuai Dai, Xunlai Chen |
CVPR | 4 |
| 2024 | Codebook Transfer with Part-of-Speech for Vector-Quantized Image ModelingabstractVector-Quantized Image Modeling (VQIM) is a fundamental research problem in image synthesis, which aims to represent an image with a discrete token sequence. Existing studies effectively address this problem by learning a discrete codebook from scratch and in a code-independent manner to quantize continuous representations into discrete tokens. However, learning a codebook from scratch and in a code-independent manner is highly challenging, which may be a key reason causing codebook collapse, i.e., some code vectors can rarely be optimized without regard to the relationship between codes and good codebook priors such that die off finally. In this paper, inspired by pretrained language models, we find that these language models have actually pretrained a superior codebook via a large number of text corpus, but such information is rarely exploited in VQIM. To this end, we propose a novel codebook transfer framework with part-of-speech, called VQCT, which aims to transfer a well-trained codebook from pretrained language models to VQIM for robust codebook learning. Specifically, we first introduce a pretrained codebook from language models and part-of-speech knowledge as priors. Then, we construct a vision-related codebook with these priors for achieving codebook transfer. Finally, a novel codebook transfer network is designed to exploit abundant semantic relationships between codes contained in pretrained codebooks for robust VQIM codebook learning. Experimental results on four datasets show that our VQCT method achieves superior VQIM performance over previous state-of-the-art methods. Baoquan Zhang, Huaibin Wang, Chuyao Luo, Xutao Li 0003, Guotao Liang, Yunming Ye, Xiaochen Qi |
CVPR | 1 |
| 2024 | LG-VQ: Language-Guided Codebook LearningabstractVector quantization (VQ) is a key technique in high-resolution and high-fidelity image synthesis, which aims to learn a codebook to encode an image with a sequence of discrete codes and then generate an image in an auto-regression manner.
Although existing methods have shown superior performance, most methods prefer to learn a single-modal codebook (\emph{e.g.}, image), resulting in suboptimal performance when the codebook is applied to multi-modal downstream tasks (\emph{e.g.}, text-to-image, image captioning) due to the existence of modal gaps.
In this paper, we propose a novel language-guided codebook learning framework, called LG-VQ, which aims to learn a codebook that can be aligned with the text to improve the performance of multi-modal downstream tasks. Specifically, we first introduce pre-trained text semantics as prior knowledge, then design two novel alignment modules (\emph{i.e.}, Semantic Alignment Module, and Relationship Alignment Module) to transfer such prior knowledge into codes for achieving codebook text alignment.
In particular, our LG-VQ method is model-agnostic, which can be easily integrated into existing VQ models. Experimental results show that our method achieves superior performance on reconstruction and various multi-modal downstream tasks. Guotao Liang, Baoquan Zhang, Yaowei Wang 0001, Yunming Ye, Xutao Li 0003, Huaibin Wang, Chuyao Luo, Kola Ye, Linfeng Luo |
NeurIPS | 2 |
| 2024 | Facilitating interaction between partial differential equation-based dynamics and unknown dynamics for regional wind speed prediction
Shidong Chen, Baoquan Zhang, Xutao Li 0001, Yunming Ye, Kenghong Lin |
Neural Networks | 2 |
| 2024 | AMANet: An Adaptive Memory Attention Network for video cloud detection
Shanshan Feng 0001, Yingling Quan, Yunming Ye, Yong Xu 0001, Xutao Li 0003, Baoquan Zhang |
Pattern Recognit. | 7 |
| 2024 | Spherical Neural Operator Network for Global Weather PredictionabstractGlobal weather forecast is an important spatial-temporal prediction problem, which can provide numerous societal benefits such as extreme weather forewarning, traffic scheduling, and agricultural planning. Though many spatial-temporal prediction models have been proposed, they suffer from two drawbacks for global weather forecasts, namely (i) ignoring the physical mechanism and spherical characteristics and (ii) not effectively exploiting the global and local correlations. To address the above drawbacks, in this paper, we formalize global weather state dynamics as partial differential equations (PDEs) in spherical space and infer the state of the global weather system by solving these PDEs. Specifically, we use Green’s function method to solve the PDEs and find that the solution of the spherical PDEs can be obtained by the spherical convolution. We further proposed a novel Spherical Neural Operator, SNO, which consists of spherical convolution and vanilla convolution. The former is used to solve these PDEs and model the global correlations in spherical space, and the latter is used to capture the local correlations. Upon the operator, a global weather prediction model is developed. Extensive experimental results demonstrate the effectiveness and superiority of our method over state-of-the-art approaches. Kenghong Lin, Xutao Li 0003, Yunming Ye, Shanshan Feng 0001, Baoquan Zhang, Guangning Xu |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | LGCNet: A Cloud Detection Method in Remote Sensing Images Using Local and Global SemanticsabstractDetecting and eliminating clouds is a crucial step in remote sensing image (RSI) preprocessing. The removal of clouds can significantly enhance the performance of subsequent remote sensing applications. Existing deep learning (DL)-based cloud detection methods extract semantic information to improve feature representation and, subsequently, detection performance. However, these methods do not fully utilize the potential of context semantic information. Besides, to capture semantics from large receptive fields, they employ convolution operators with large kernel sizes, which results in high computational costs. Thus, these computationally heavy models are not suitable for resource-limited devices, particularly satellites. To address this issue, we propose a cloud detection model, LGCNet. This model efficiently extracts both local and global contextual information, fully utilizing semantics while reducing resource usage. LGCNet is built on an encoder-decoder structure. Specifically, the encoder extracts local scale-aware semantics through proposed local semantic blocks (LSBs), which are then skip-connected to the decoder. This approach provides adaptive and diverse local contextual information. On the top of the encoder, the high-level global semantics are captured via the proposed global feature TransBlock (GFTB). A variety of extracted semantics ensure improved detection performance. We evaluate the proposed method using two public datasets: LandSat8 and Moderate-Resolution Imaging Spectroradiometer (MODIS). We conducted experiments on both a server and an edge computing device. Our extensive experiments revealed that LGCNet outperforms other lightweight cloud detection and semantic segmentation methods in terms of performance and computational load. Shanshan Feng 0001, Huaibin Wang, Baoquan Zhang, Pengjuan Yao, Chuyao Luo, Yunming Ye, Yong Xu 0001, Xutao Li 0003 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | A Practical Online Incremental Learning Framework for Precipitation NowcastingabstractPrecipitation nowcasting plays an important role in our life. Many deep learning-based methods are proposed for precipitation nowcasting by predicting radar echo sequence over the past years, and achieving better performance than traditional approaches. However, all of them are based on a static model, which is trained in offline learning and does not adapt to real-time changing precipitation data. Recently, online incremental learning (OIL) has been proposed to dynamically update the model by continually learning new data and preventing the forgetting of historical knowledge in an online fashion. While effective, existing OIL approaches that focus on a classification task are not suitable for the regression task of precipitation nowcasting. To fill this gap, we try to propose a novel OIL framework for precipitation nowcasting. By analyzing its characteristics, we find three challenges: 1) the distributions of radar echo maps in different rainfall events are different; 2) in each rainfall event, there is always an inevitable delay between the timestamps of the training and testing samples; and 3) the real-time requirement for model prediction is very high, which has strict limitations on the training speed of the model. Based on these observations, we propose a practical OIL framework based on gradient activation mapping (GAM). It can be mainly divided into three components: 1) recall training strategy (RTS) is used to eliminate the interference caused by distributions of different events; 2) iterative approximation training (IAT) is designed to align the timestamps of the training and testing; and 3) moreover, we propose gradient activation mapping weight (GAMW) to improve the training effectiveness. Extensive experiments show that the proposed framework can improve the performance of the model stably and effectively. Especially in those heavy rainfall regions where it usually causes more threat to human activity, the improvement is significant. Chuyao Luo, Zheng Zhang 0046, Huiwei Lin, Baoquan Zhang, Xutao Li 0003, Yunming Ye |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | MCSDNet: Mesoscale Convective System Detection Network via Multiscale Spatiotemporal InformationabstractThe accurate detection of mesoscale convective systems (MCSs) is crucial for meteorological monitoring due to their potential to cause significant destruction through severe weather phenomena, such as hail, thunderstorms, and heavy rainfall. However, the existing methods for MCS detection mostly targets on single-frame detection, which just considers the static characteristics and ignores the temporal evolution in the life cycle of MCS. In this article, we propose a novel encoder-decoder neural network named mesoscale convective system detection network (MCSDNet) to detect MCS regions. MCSDNet has a simple architecture and is easy to expand. Different from the previous models, MCSDNet targets on multiframes detection and leverages multiscale spatiotemporal information in remote sensing imagery (RSI). As far as we know, it is the first work to utilize multiscale spatiotemporal information to detect MCS regions. First, we design a multiscale spatiotemporal information module to extract multilevel semantic from different encoder levels, which makes our models can extract more detail spatiotemporal features. Second, spatiotemporal mix unit (STMU), a dual spatiotemporal attention, is introduced to MCSDNet to capture both intraframe features and interframe. Finally, we present MCS remote sensing image (MCSRSI) the first publicly available dataset for multiframes MCS detection based on FY-4A satellite. We also conduct several experiments on MCSRSI and find that our proposed MCSDNet achieves the best performance on MCS detection task when comparing with other baseline methods. We hope that the combination of our open-access dataset and promising results will encourage the future research for MCS detection task and provide a robust framework for related tasks in atmospheric science. Our code is available at:https://github.com/250HandsomeLiang/MCSDNet.git Baoquan Zhang, Jiajun Liang, Rui Ye 0002, Chuyao Luo, Xutao Li 0003, Yunming Ye, Xukai Fu |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | PCR: Proxy-Based Contrastive Replay for Online Class-Incremental Continual LearningabstractOnline class-incremental continual learning is a specific task of continual learning. It aims to continuously learn new classes from data stream and the samples of data stream are seen only once, which suffers from the catastrophic forgetting issue, i.e., forgetting historical knowledge of old classes. Existing replay-based methods effectively alleviate this issue by saving and replaying part of old data in a proxy-based or contrastive-based replay manner. Although these two replay manners are effective, the former would incline to new classes due to class imbalance issues, and the latter is unstable and hard to converge because of the limited number of samples. In this paper, we conduct a comprehensive analysis of these two replay manners and find that they can be complementary. Inspired by this finding, we propose a novel replay-based method called proxy-based contrastive replay (PCR). The key operation is to replace the contrastive samples of anchors with corresponding proxies in the contrastive-based way. It alleviates the phenomenon of catastrophic forgetting by effectively addressing the imbalance issue, as well as keeps a faster convergence of the model. We conduct extensive experiments on three real-world benchmark datasets, and empirical results consistently demonstrate the superiority of PCR over various state-of-the-art methods11https://github.com/FelixHuiweiLin/PCR. Huiwei Lin, Baoquan Zhang, Shanshan Feng 0001, Xutao Li 0001, Yunming Ye |
CVPR | 2 |
| 2023 | UER: A Heuristic Bias Addressing Approach for Online Continual LearningabstractOnline continual learning aims to continuously train neural networks from a continuous data stream with a single pass-through data. As the most effective approach, the rehearsal-based methods replay part of previous data. Commonly used predictors in existing methods tend to generate biased dot-product logits that prefer to the classes of current data, which is known as a bias issue and a phenomenon of forgetting. Many approaches have been proposed to overcome the forgetting problem by correcting the bias; however, they still need to be improved in online fashion. In this paper, we try to address the bias issue by a more straightforward and more efficient method. By decomposing the dot-product logits into an angle factor and a norm factor, we empirically find that the bias problem mainly occurs in the angle factor, which can be used to learn novel knowledge as cosine logits. On the contrary, the norm factor abandoned by existing methods helps remember historical knowledge. Based on this observation, we intuitively propose to leverage the norm factor to balance the new and old knowledge for addressing the bias. To this end, we develop a heuristic approach called unbias experience replay (UER). UER learns current samples only by the angle factor and further replays previous samples by both the norm and angle factors. Extensive experiments on three datasets show that UER achieves superior performance over various state-of-the-art methods. The code is in https://github.com/FelixHuiweiLin/UER. Huiwei Lin, Shanshan Feng 0001, Baoquan Zhang, Hongliang Qiao, Xutao Li 0003, Yunming Ye |
ACM Multimedia | 3 |
| 2023 | Multi-view knowledge graph fusion via knowledge-aware attentional graph neural network
Zhichao Huang 0001, Xutao Li 0003, Yunming Ye, Baoquan Zhang, Guangning Xu, Wensheng Gan |
Appl. Intell. | 4 |
| 2023 | WDMNet: Modeling diverse variations of regional wind speed for multi-step predictions
Rui Ye 0002, Shanshan Feng 0001, Xutao Li 0003, Yunming Ye, Baoquan Zhang, Yaowei Wang 0001 |
Neural Networks | 5 |
| 2023 | PEPNet: A barotropic primitive equations-based network for wind speed prediction
Rui Ye 0002, Baoquan Zhang, Xutao Li 0003, Yunming Ye |
Neural Networks | 2 |
| 2023 | Prototype Completion for Few-Shot LearningabstractFew-shot learning (FSL) aims to recognize novel classes with few examples. Pre-training based methods effectively tackle the problem by pre-training a feature extractor and then fine-tuning it through the nearest centroid based meta-learning. However, results show that the fine-tuning step makes marginal improvements. In this paper, 1) we figure out the reason, i.e., in the pre-trained feature space, the base classes already form compact clusters while novel classes spread as groups with large variances, which implies that fine-tuning feature extractor is less meaningful; 2) instead of fine-tuning feature extractor, we focus on estimating more representative prototypes. Consequently, we propose a novel prototype completion based meta-learning framework. This framework first introduces primitive knowledge (i.e., class-level part or attribute annotations) and extracts representative features for seen attributes as priors. Second, a part/attribute transfer network is designed to learn to infer the representative features for unseen attributes as supplementary priors. Finally, a prototype completion network is devised to learn to complete prototypes with these priors. Moreover, to avoid the prototype completion error, we further develop a Gaussian based prototype fusion strategy that fuses the mean-based and completed prototypes by exploiting the unlabeled samples. At last, we also develop an economic prototype completion version for FSL, which does not need to collect primitive knowledge, for a fair comparison with existing FSL methods without external knowledge. Extensive experiments show that our method: i) obtains more accurate prototypes; ii) achieves superior performance on both inductive and transductive FSL settings. Baoquan Zhang, Xutao Li 0003, Yunming Ye, Shanshan Feng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Knowledge-enhanced Prompt-tuning for Stance DetectionabstractInvestigating public attitudes on social media is important in opinion mining systems. Stance detection aims to analyze the attitude of an opinionated text (e.g., favor, neutral, or against) toward a given target. Existing methods mainly address this problem from the perspective of fine-tuning. Recently, prompt-tuning has achieved success in natural language processing tasks. However, conducting prompt-tuning methods for stance detection in real-world remains a challenge for several reasons: (1) The text form of stance detection is usually short and informal, which makes it difficult to design label words for the verbalizer. (2) The tweet text may not explicitly give the attitude. Instead, users may use various hashtags or background knowledge to express stance-aware perspectives. In this article, we first propose a prompt-tuning-based framework that performs stance detection in a cloze question manner. Specifically, a knowledge-enhanced prompt-tuning framework (KEprompt) method is designed, which consists of an automatic verbalizer (AutoV) and background knowledge injection (BKI). Specifically, in AutoV, we introduce a semantic graph to build a better mapping from the predicted word of the pretrained language model and detection labels. In BKI, we first propose a topic model for learning hashtag representation and introduce ConceptGraph as the supplement of the target. At last, we present a challenging dataset for stance detection, where all stance categories are expressed in an implicit manner. Extensive experiments on a large real-world dataset demonstrate the superiority of KEprompt over state-of-the-art methods. Hu Huang 0009, Bowen Zhang 0005, Xiang-Yang Li 0001, Baoquan Zhang, Yuxi Sun 0002, Chuyao Luo, Cheng Peng 0003 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |
| 2023 | PMDB: A Range-Based Key-Value Store on Hybrid NVM-Storage SystemsabstractEmerging Nov-Volatile Memory (NVM) may replace DRAM as main memory in future computers. However, data will likely still be stored on storage due to the enormous large size of available data. We investigate how key-value stores can be efficiently designed and implemented in a hybrid system, called NVM-Storage system, consisting of NVM as memory and traditional storage. We first discuss the performance trade-offs among Put, Get, and Range Query of the existing designs. Then, we propose PMDB, a range-based key-value store on NVM-Storage systems. PMDB achieves good performance for Put, Get and Range Query at the same time by utilizing a range-based data management and deploying a light-weight index on NVM. We compare PMDB with the state-of-the-art schemes including SLM-DB [21] and MatrixKV [40] for hybrid NVM-storage systems. Evaluation results indicate that in workloads with mixed Put, Get and Range Queries, PMDB outperforms existing key-value stores by$1.16\times$–$2.49\times$. Baoquan Zhang, Haoyu Gong, David Hung-Chang Du |
IEEE Trans. Computers | 1 |
| 2023 | Adaptive Transfer of Graph Neural Networks for Few-Shot Molecular Property PredictionabstractFew-Shot Molecular Property Prediction (FSMPP) is an improtant task on drug discovery, which aims to learn transferable knowledge from base property prediction tasks with sufficient data for predicting novel properties with few labeled molecules. Its key challenge is how to alleviate the data scarcity issue of novel properties. Pretrained Graph Neural Network (GNN) based FSMPP methods effectively address the challenge by pre-training a GNN from large-scale self-supervised tasks and then finetuning it on base property prediction tasks to perform novel property prediction. However, in this paper, we find that the GNN finetuning step is not always effective, which even degrades the performance of pretrained GNN on some novel properties. This is because these molecule-property relationships among molecules change across different properties, which results in the finetuned GNN overfits to base properties and harms the transferability performance of pretrained GNN on novel properties. To address this issue, in this paper, we propose a novel Adaptive Transfer framework of GNN for FSMPP, called ATGNN, which transfers the knowledge of pretrained and finetuned GNNs in a task-adaptive manner to adapt novel properties. Specifically, we first regard the pretrained and finetuned GNNs as model priors of target-property GNN. Then, a task-adaptive weight prediction network is designed to leverage these priors to predict target GNN weights for novel properties. Finally, we combine our ATGNN framework with existing FSMPP methods for FSMPP. Extensive experiments on four real-world datasets, i.e., Tox21, SIDER, MUV, and ToxCast, show the effectiveness of our ATGNN framework. Baoquan Zhang, Chuyao Luo, Hao Jiang 0051, Shanshan Feng 0001, Xutao Li 0003, Bowen Zhang 0005, Yunming Ye |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2023 | MetaDT: Meta Decision Tree With Class Hierarchy for Interpretable Few-Shot LearningabstractFew-Shot Learning (FSL) is a challenging task, which aims to recognize novel classes with few examples. Recently, lots of methods have been proposed from the perspective of meta-learning and representation learning. However, few works focus on the interpretability of FSL decision process. In this paper, we take a step towards the interpretable FSL by proposing a novel meta-learning based decision tree framework, namely, MetaDT. In particular, the FSL interpretability is achieved from two aspects, i.e., a concept aspect and a visual aspect. On the concept aspect, we first introduce a tree-like concept hierarchy as FSL prior. Then, resorting to the prior, we split each few-shot task to a set of subtasks with different concept levels and then perform class prediction via a model of decision tree. The advantage of such design is that a sequence of high-level concept decisions that lead up to a final class prediction can be obtained, which clarifies the FSL decision process. On the visual aspect, a set of subtask-specific classifiers with visual attention mechanism is designed to perform decision at each node of the decision tree. As a result, a subtask-specific heatmap visualization can be obtained to achieve the decision interpretability of each tree node. At last, to alleviate the data scarcity issue of FSL, we regard the prior of concept hierarchy as an undirected graph, and then design a graph convolution-based decision tree inference network as our meta-learner to infer parameters of the decision tree. Extensive experiments on performance comparison and interpretability analysis show superiority of our MetaDT. Baoquan Zhang, Hao Jiang 0051, Xutao Li 0003, Shanshan Feng 0001, Yunming Ye, Rui Ye 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | TRCDNet: A Transformer Network for Video Cloud DetectionabstractIn Remote Sensing Image (RSI) pre-processing steps, detecting and removing cloudy areas is a critical task. Recently, cloud detection methods based on deep neural networks achieve outstanding performance over traditional methods. Current approaches mostly focus on cloud detection on a single image captured by polar-orbiting satellites. However, there is another type of meteorological satellite - geostationary satellite, which can capture temporal consecutive frames of a particular location. Therefore, the cloud detection task targeting at geostationary satellite can be treated as a video cloud detection task. And in addition to extracting features on a single image, extracting and making full use of the relations between sequential frames is also important. To tackle this problem, we design a deep learning video cloud detection model: Transformer Network for Video Cloud Detection (TRCDNet). The proposed network is based on the encoder-decoder structure. In the encoder, the module ContextGhostLayer is proposed to encode more semantic information to tackle the challenging problems like thin cloud in RSIs. Besides, we design a transformer-based Video Sequence Transformer (VSTR) block. Based on attention mechanism, VSTR can fully extract the across-frame relations. In the proposed decoder, the cloud masks are recovered gradually to the same scale as the input image. To evaluate the methods, we create a Video Cloud Detection dataset based on the captured videos from Fengyun 4 (FY-4) satellite: Fengyun4aCloud. Extensive experiments of current cloud detection methods, semantic segmentation methods, and video semantic segmentation methods indicate that the designed TRCDNet achieves state-of-art performance in video cloud detection. Shanshan Feng 0001, Yingling Quan, Yunming Ye, Xutao Li 0003, Yong Xu 0001, Baoquan Zhang, Zhihao Chen 0010 |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2022 | MetaNODE: Prototype Optimization as a Neural ODE for Few-Shot LearningabstractFew-Shot Learning (FSL) is a challenging task, i.e., how to recognize novel classes with few examples? Pre-training based methods effectively tackle the problem by pre-training a feature extractor and then predicting novel classes via a cosine nearest neighbor classifier with mean-based prototypes. Nevertheless, due to the data scarcity, the mean-based prototypes are usually biased. In this paper, we attempt to diminish the prototype bias by regarding it as a prototype optimization problem. To this end, we propose a novel meta-learning based prototype optimization framework to rectify prototypes, i.e., introducing a meta-optimizer to optimize prototypes. Although the existing meta-optimizers can also be adapted to our framework, they all overlook a crucial gradient bias issue, i.e., the mean-based gradient estimation is also biased on sparse data. To address the issue, we regard the gradient and its flow as meta-knowledge and then propose a novel Neural Ordinary Differential Equation (ODE)-based meta-optimizer to polish prototypes, called MetaNODE. In this meta-optimizer, we first view the mean-based prototypes as initial prototypes, and then model the process of prototype optimization as continuous-time dynamics specified by a Neural ODE. A gradient flow inference network is carefully designed to learn to estimate the continuous gradient flow for prototype dynamics. Finally, the optimal prototypes can be obtained by solving the Neural ODE. Extensive experiments on miniImagenet, tieredImagenet, and CUB-200-2011 show the effectiveness of our method. Baoquan Zhang, Xutao Li 0003, Shanshan Feng 0001, Yunming Ye, Rui Ye 0002 |
AAAI | 1 |
| 2022 | Sentiment Interpretable Logic Tensor Network for Aspect-Term Sentiment AnalysisabstractAspect-term sentiment analysis (ATSA) is an important task that aims to infer the sentiment towards the given aspect-terms. It is often required in the industry that ATSA should be performed with interpretability, computational efficiency and high accuracy. However, such an ATSA method has not yet been developed. This study aims to develop an ATSA method that fulfills all these requirements. To achieve the goal, we propose a novel Sentiment Interpretable Logic Tensor Network (SILTN). SILTN is interpretable because it is a neurosymbolic formalism and a computational model that supports learning and reasoning about data with a differentiable first-order logic language (FOL). To realize SILTN with high inferring accuracy, we propose a novel learning strategy called the two-stage syntax knowledge distillation (TSynKD). Using widely used datasets, we experimentally demonstrate that the proposed TSynKD is effective for improving the accuracy of SILTN, and the SILTN has both high interpretability and computational efficiency. Bowen Zhang 0005, Zhichao Huang 0001, Hu Huang 0009, Baoquan Zhang, Xianghua Fu, Liwen Jing 0001 |
COLING | 5 |
| 2022 | Hyperbolic Knowledge Transfer with Class Hierarchy for Few-Shot LearningabstractFew-shot learning (FSL) aims to recognize a novel class with very few instances, which is a challenging task since it suffers from a data scarcity issue. One way to effectively alleviate this issue is introducing explicit knowledge summarized from human past experiences to achieve knowledge transfer for FSL. Based on this idea, in this paper, we introduce the explicit knowledge of class hierarchy (i.e., the hierarchy relations between classes) as FSL priors and propose a novel hyperbolic knowledge transfer framework for FSL, namely, HyperKT. Our insight is, in the hyperbolic space, the hierarchy relation between classes can be well preserved by resorting to the exponential growth characters of hyperbolic volume, so that better knowledge transfer can be achieved for FSL. Specifically, we first regard the class hierarchy as a tree-like structure. Then, 1) a hyperbolic representation learning module and a hyperbolic prototype inference module are employed to encode/infer each image and class prototype to the hyperbolic space, respectively; and 2) a novel hierarchical classification and relation reconstruction loss are carefully designed to learn the class hierarchy. Finally, the novel class prediction is performed in a nearest-prototype manner. Extensive experiments on three datasets show our method achieves superior performance over state-of-the-art methods, especially on 1-shot tasks. Baoquan Zhang, Hao Jiang 0051, Shanshan Feng 0001, Xutao Li 0003, Yunming Ye, Rui Ye 0002 |
IJCAI | 1 |
| 2022 | SPLNet: A sequence-to-one learning network with time-variant structure for regional wind speed prediction
Rui Ye 0002, Shanshan Feng 0001, Xutao Li 0003, Yunming Ye, Baoquan Zhang, Chuyao Luo |
Inf. Sci. | 5 |
| 2022 | DynamicNet: A time-variant ODE network for multi-step wind speed prediction
Rui Ye 0002, Xutao Li 0003, Yunming Ye, Baoquan Zhang |
Neural Networks | 4 |
| 2022 | ECDNet: A bilateral lightweight cloud detection network for remote sensing images
Shanshan Feng 0001, Xutao Li 0001, Yunming Ye, Baoquan Zhang, Zhihao Chen 0010, Yingling Quan |
Pattern Recognit. | 5 |
| 2022 | LWCDnet: A Lightweight Network for Efficient Cloud Detection in Remote Sensing ImagesabstractCloud detection is the task of detecting cloud areas in remote sensing images, and it has attracted extensive research interest. Recently, deep learning-based methods have been proposed and achieved great performance for cloud detection. However, due to the satellite’s limitation in storage and memory, existing deep learning approaches, which suffer from extensive computation and large model size, are almost impossible to be deployed on satellites. To fill this gap, we target at studying effective and efficient cloud detection solutions that are suitable for satellites. In this paper, we develop a lightweight autoencoder-based cloud detection method, namely LWCDnet. In the encoder part, the designed novel lightweight dual-branch block (LWDBB) in the backbone extracts spatial and contextual information concurrently. Moreover, a lightweight feature pyramid module (LWFPM) is proposed to capture high-level multi-scale contextual information. In the decoder part, the lightweight feature fusion module (LWFFM) compensates for the missing spatial and detail information from the encoder to the high-level feature maps. We evaluate the proposed method on two public datasets: LandSat8 and MODIS. Extensive experiments demonstrate that the proposed LWCDnet achieves comparable accuracy as the-state-of-art cloud detection methods and lightweight semantic segmentation algorithms. Meantime LWCDnet has much less computation burden with smaller model size. Shanshan Feng 0001, Xiaofei Yang 0002, Yunming Ye, Xutao Li 0003, Baoquan Zhang, Zhihao Chen 0010, Yingling Quan |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2022 | SGMNet: Scene Graph Matching Network for Few-Shot Remote Sensing Scene ClassificationabstractFew-Shot Remote Sensing Scene Classification (FSRSSC) is an important task, which aims to recognize novel scene classes with few examples. Recently, several studies attempt to address the FSRSSC problem by following few-shot natural image classification methods. These existing methods have made promising progress and achieved superior performance. However, they all overlook two unique characteristics of remote sensing images: (i)object co-occurrencethat multiple objects tend to appear together in a scene image and (ii)object spatial correlationthat these co-occurrence objects are distributed in the scene image following some spatial structure patterns. Such unique characteristics are very beneficial for FSRSSC, which can effectively alleviate the scarcity issue of labeled remote sensing images since they can provide more refined descriptions for each scene class. To fully exploit these characteristics, we propose a novel scene graph matching-based meta-learning framework for FSRSSC, called SGMNet. In this framework, a scene graph construction module is carefully designed to represent each test remote sensing image or each scene class as a scene graph, where the nodes reflect these co-occurrence objects meanwhile the edges capture the spatial correlations between these co-occurrence objects. Then, a scene graph matching module is further developed to evaluate the similarity score between each test remote sensing image and each scene class. Finally, based on the similarity scores, we perform the scene class prediction via a nearest neighbor classifier. We conduct extensive experiments on UCMerced LandUse, WHU19, AID, and NWPU-RESISC45 datasets. The experimental results show that our method obtains superior performance over the previous state-of-the-art methods. Baoquan Zhang, Shanshan Feng 0001, Xutao Li 0003, Yunming Ye, Rui Ye 0002, Hao Jiang 0051 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2021 | Prototype Completion With Primitive Knowledge for Few-Shot LearningabstractFew-shot learning is a challenging task, which aims to learn a classifier for novel classes with few examples. Pre-training based meta-learning methods effectively tackle the problem by pre-training a feature extractor and then fine-tuning it through the nearest centroid based meta-learning. However, results show that the fine-tuning step makes very marginal improvements. In this paper, 1) we figure out the key reason, i.e., in the pre-trained feature space, the base classes already form compact clusters while novel classes spread as groups with large variances, which implies that fine-tuning the feature extractor is less meaningful; 2) instead of fine-tuning the feature extractor, we focus on estimating more representative prototypes during meta-learning. Consequently, we propose a novel prototype completion based meta-learning framework. This framework first introduces primitive knowledge (i.e., class-level part or attribute annotations) and extracts representative attribute features as priors. Then, we design a prototype completion network to learn to complete prototypes with these priors. To avoid the prototype completion error caused by primitive knowledge noises or class differences, we further develop a Gaussian based prototype fusion strategy that combines the mean-based and completed prototypes by exploiting the unlabeled samples. Extensive experiments show that our method: (i) can obtain more accurate prototypes; (ii) out-performs state-of-the-art techniques by 2%~9% in terms of classification accuracy. Our code is available online1. Baoquan Zhang, Xutao Li 0003, Yunming Ye, Zhichao Huang 0001, Lisai Zhang |
CVPR | 1 |
| 2021 | Learn to abstract via concept graph for weakly-supervised few-shot learning
Baoquan Zhang, Ka-Cheong Leung, Xutao Li 0001, Yunming Ye |
Pattern Recognit. | 1 |
| 2021 | TrackLace: Data Management for Interlaced Magnetic RecordingabstractInterlaced Magnetic Recording (IMR) is a promising technology which achieves higher data density and lower write amplification (WA) than Shingled Magnetic Recording (SMR). In IMR, top tracks and bottom tracks are interlaced so each bottom track is partially overlapped with two adjacent top tracks. Top tracks can be updated without any WA, but bottom track updates require reading and rewriting of affected valid data on the two neighboring top tracks. There are few published studies discussing WA in IMR drives. We propose TrackLace to reduce WA for IMR. TrackLace consists of three techniques: Z-Alloc allocates user data to the tracks in alternating directions and spreads unallocated tracks among allocated tracks; Top-Buffer opportunistically utilizes unallocated top tracks to buffer bottom track updates; and Block-Swap progressively swaps bottom track hot data with top track cold data during high space utilization. To further optimize TrackLace performance, we propose a virtual frame design that can keep the relocated block (due to Top-Buffer or Block-Swap) close to its original location and an adaptive buffering mechanism that can avoid unnecessary redirections depending on the write locality. Evaluations show that TrackLace can reduce WA by 45 percent and lower average latency by 31percent compared with baseline schemes. Fenggang Wu, Bingzhe Li, Baoquan Zhang, Zhichao Cao 0002, Jim Diehl, Hao Wen 0001, David Hung-Chang Du |
IEEE Trans. Computers | 3 |
| 2021 | NVLSM: A Persistent Memory Key-Value Store Using Log-Structured Merge Tree with Accumulative CompactionabstractComputer systems utilizing byte-addressable Non-Volatile Memory ( NVM ) as memory/storage can provide low-latency data persistence. The widely used key-value stores using Log-Structured Merge Tree ( LSM-Tree ) are still beneficial for NVM systems in aspects of the space and write efficiency. However, the significant write amplification introduced by the leveled compaction of LSM-Tree degrades the write performance of the key-value store and shortens the lifetime of the NVM devices. The existing studies propose new compaction methods to reduce write amplification. Unfortunately, they result in a relatively large read amplification. In this article, we propose NVLSM, a key-value store for NVM systems using LSM-Tree with new accumulative compaction. By fully utilizing the byte-addressability of NVM, accumulative compaction uses pointers to accumulate data into multiple floors in a logically sorted run to reduce the number of compactions required. We have also proposed a cascading searching scheme for reads among the multiple floors to reduce read amplification. Therefore, NVLSM reduces write amplification with small increases in read amplification. We compare NVLSM with key-value stores using LSM-Tree with two other compaction methods: leveled compaction and fragmented compaction. Our evaluations show that NVLSM reduces write amplification by up to 67% compared with LSM-Tree using leveled compaction without significantly increasing the read amplification. In write-intensive workloads, NVLSM reduces the average latency by 15.73%–41.2% compared to other key-value stores. Baoquan Zhang, David Hung-Chang Du |
ACM Trans. Storage | 1 |
| 2020 | AC-Key: Adaptive Caching for LSM-based Key-Value Stores
Fenggang Wu, Ming-Hong Yang, Baoquan Zhang, David Hung-Chang Du |
USENIX ATC | 3 |
| 2020 | Idler : I/O Workload Controlling for Better Responsiveness on Host-Aware Shingled Magnetic Recording DrivesabstractHost-Aware/Drive-Managed Shingled Magnetic Recording (SMR) drives can accept non-sequential writes using a buffer called media cache. Data in the media cache will be migrated to its designated location by a cleaning process if the buffer is full (blocking cleaning) or the drive is idle (idle cleaning). However, blocking cleanings can severely extend the I/O response time. Therefore, it is crucial to fully understand the cleaning process and find ways of mitigating the caused performance degradation. In this article we further evaluate the cleaning process and propose a potential remedy scheme called Idler on Host-Aware SMR drives. Idler adaptively induces idle cleanings based on dynamic workload characteristics and media cache usages to reduce the severity of blocking cleanings. Our evaluations show that in the workloads with a small non-sequential write ratio (about 10 percent), Idler can reduce the tail response time and the workload finish time by 56-88 and 10-23 percent, respectively, compared with those without such control. With the help of an external write buffer on an SSD, the tail response time of SMR drives with Idler can be closer to that of conventional disk drives. Baoquan Zhang, Ming-Hong Yang, Xuchao Xie, David Hung-Chang Du |
IEEE Trans. Computers | 1 |
| 2019 | ZoneAlloy: Elastic Data and Space Management for Hybrid SMR Drives
Fenggang Wu, Bingzhe Li, Zhichao Cao 0002, Baoquan Zhang, Ming-Hong Yang, Hao Wen 0001, David Hung-Chang Du |
HotStorage | 4 |
| 2019 | The Dual-Aspect Geometric Terrain Correction Method Using GF-3 Satellite DataabstractThe GF-3 satellite is the first full-polarization SAR satellite in China, which operates in the C band with a resolution of 1m. This paper presents a dual-aspect geometric terrain correction method to overcome the inherent shortages of SAR image such as foreshortening, shadow and layover. The results using GF-3 satellite data show that the method can effectively eliminate layover and shadow distortions in SAR images. This method solves the geometric correction problem that cannot be solved with a single SAR image. Xiaolan Qiu, Baoquan Zhang, Feng Wang 0019 |
IGARSS | 3 |
| 2018 | Improving Data Integrity in Linux Software RAID with Protection Information (T10-PI)abstractThe T10 DIF (Data Integrity Field) and DIX (Data Integrity Extension) specifications provide mechanisms to guarantee end-to-end data integrity and protection in the face of silent data corruption in modern storage systems. However, the Multiple Devices (MD) software RAID driver in Linux does not fully leverage these capabilities to provide such end-to-end guarantees with widely-used RAID modes such as 5 and 6, thereby causing an "integrity gap" in the Linux I/O stack. This paper describes the design and performance characteristics of a DIX-aware MD module that plugs this integrity gap with minimal overhead to client applications. A PI (Protection Information) operator is added in MD to handle the PI-related operations, and dedicated buffers for PI are allocated and managed in MD RAID-5/6 personality's stripe structures to generate, store, and verify the PI. This allows seamless exchange of PI information among end-applications running in user mode, file systems, the linux block layer, and PI-capable HBAs and drives. Our evaluations show that the DIX-aware MD module has the capability of detecting SDC with the tolerable performance penalty. Baoquan Zhang, Raghunath Rajachandrasekar, Lance Evans, David Hung-Chang Du |
CCGrid | 1 |
| 2018 | Data Management Design for Interlaced Magnetic Recording
Fenggang Wu, Baoquan Zhang, Zhichao Cao 0002, Hao Wen 0001, Bingzhe Li, Jim Diehl, David Hung-Chang Du |
HotStorage | 2 |
| 2017 | IPRM: IP core resource multiplexing of core wrapper design for reducing test application time in DVFS-based multicore SoCs
Libao Deng, Baoquan Zhang, Chengyu Jin |
Integr. | 2 |
| 2017 | Performance Evaluation of Host Aware Shingled Magnetic Recording (HA-SMR) DrivesabstractShingled Magnetic Recording (SMR) drives can benefit large-scale storage systems by reducing the Total Cost of Ownership (TCO) of dealing with explosive data growth. Among all existing SMR models, Host Aware SMR (HA-SMR) looks the most promising for its backward compatibility with legacy I/O stacks and its ability to use new SMR-specific APIs to support host I/O stack optimization. Building storage systems using HA-SMR drives calls for a deep understanding of the drive's performance characteristics. To accomplish this, we conduct in-depth performance evaluations on HA-SMR drives with a special emphasis on the performance implications of the SMR-specific APIs and how these drives can be deployed in large storage systems. We discover both favorable and adverse effects of using HA-SMR drives under various workloads. We also investigate the drive's performance under legacy production environments using real-world enterprise traces. Finally, we propose a novel host-controlled buffer that can help to reduce the severity of the decline in HA-SMR performance under our discovered unfavorable I/O access patterns. Without a detailed comprehensive design, we show the potential of the host-controlled buffer by a case study. Fenggang Wu, Ziqi Fan, Ming-Chang Yang, Baoquan Zhang, Xiongzi Ge, David Hung-Chang Du |
IEEE Trans. Computers | 4 |
| 2016 | Evaluating Host Aware SMR Drives
Fenggang Wu, Ming-Chang Yang, Ziqi Fan, Baoquan Zhang, Xiongzi Ge, David Hung-Chang Du |
HotStorage | 4 |