VLDB 2026 Research / reviewers in the wild / expert
Xiaoyang Qu
dblp:168/2623
· DBLP profile ↗
76ranked-venue papers
5as first author
64since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 42 · 2 first-author · 40 since 2021Graphics, computer vision, multimedia, augmented reality and games · 26 · 2 first-author · 24 since 2021Systems, architecture and hardware · 13 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Computer networks · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Vista: Scene-Aware Optimization for Streaming Video Question Answering Under Post-Hoc QueriesabstractStreaming video question answering (Streaming Video QA) poses distinct challenges for multimodal large language models (MLLMs), as video frames arrive sequentially and user queries can be issued at arbitrary timepoints. Existing solutions relying on fixed-size memory or naive compression often suffer from context loss or memory overflow, limiting their effectiveness in long-form, real-time scenarios.We present Vista, a novel framework for scene-aware streaming video QA that enables efficient and scalable reasoning over continuous video streams. The innovation of Vista can be summarized in three aspects: (1) Scene-aware segmentation. Vista dynamically clusters incoming frames into temporally and visually coherent scene units. (2) Scene-aware compression. Each scene is compressed into a compact token representation and stored in GPU memory for efficient index-based retrieval, while the full-resolution frames are offloaded to CPU memory. (3) Scene-aware recall. Upon receiving a question, relevant scenes are selectively recalled and reintegrated into the model’s input space, enabling both efficiency and completeness. Vista is model-agnostic and integrates seamlessly with a variety of vision-language backbones, enabling long-context reasoning without compromising latency or memory efficiency. Extensive experiments on StreamingBench demonstrate that Vista achieves state-of-the-art performance, establishing a strong baseline for real-world streaming video understanding. Haocheng Lu, Xiaoyang Qu, Guokuan Li, Jiguang Wan 0001, Jianzong Wang |
AAAI | 4 |
| 2026 | From Inheritance to Saturation: Disentangling the Evolution of Visual Redundancy for Architecture-Aware MLLM Inference AccelerationabstractHigh-resolution Multimodal Large Language Models (MLLMs) face prohibitive computational costs during inference due to the explosion of visual tokens. Existing acceleration strategies, such as token pruning or layer sparsity, suffer from severe “backbone dependency”, performing well on Vicuna or Mistral architectures (e.g., LLaVA) but causing significant performance degradation when transferred to architectures like Qwen. To address this, we leverage truncated matrix entropy to uncover a universal three-stage inference lifecycle, decoupling visual redundancy into universal Intrinsic Visual Redundancy (IVR) and architecture-dependent Secondary Saturation Redundancy (SSR). Guided by this insight, we propose HalfV, a framework that first mitigates IVR via a unified pruning strategy and then adaptively handles SSR based on its specific manifestation. Experiments demonstrate that HalfV achieves superior efficiency-performance trade-offs across diverse backbones. Notably, on Qwen25-VL, it retains 96.8% performance at a 4.1\times FLOPs speedup, significantly outperforming state-of-the-art baselines. Our code is available at https://github.com/civilizwa/HalfV. Xulong Zhang 0001, Yuechan Li, Xiaoyang Qu, Jianzong Wang |
ACL (1) | 4 |
| 2025 | ACCon: Angle-Compensated Contrastive Regularizer for Deep RegressionabstractIn deep regression, capturing the relationship among continuous labels in feature space is a fundamental challenge that has attracted increasing interest. Addressing this issue can prevent models from converging to suboptimal solutions across various regression tasks, leading to improved performance, especially for imbalanced regression and under limited sample sizes. However, existing approaches often rely on order-aware representation learning or distance-based weighting. In this paper, we hypothesize a linear negative correlation between label distances and representation similarities in regression tasks. To implement this, we propose an angle-compensated contrastive regularizer for deep regression, which adjusts the cosine distance between anchor and negative samples within the contrastive learning framework. Our method offers a plug-and-play compatible solution that extends most existing contrastive learning methods for regression tasks. Extensive experiments and theoretical analysis demonstrate that our proposed angle-compensated contrastive regularizer not only achieves competitive regression performance but also excels in data efficiency and effectiveness on imbalanced datasets. Botao Zhao 0001, Xiaoyang Qu, Zuheng Kang, Junqing Peng, Jing Xiao 0006, Jianzong Wang |
AAAI | 2 |
| 2025 | RUNA: Object-Level Out-of-Distribution Detection via Regional Uncertainty Alignment of Multimodal RepresentationsabstractEnabling object detectors to recognize out-of-distribution (OOD) objects is vital for building reliable systems. A primary obstacle stems from the fact that models frequently do not receive supervisory signals from unfamiliar data, leading to overly confident predictions regarding OOD objects. Despite previous progress that estimates OOD uncertainty based on the detection model and in-distribution (ID) samples, we explore using pre-trained vision-language representations for object-level OOD detection. We first discuss the limitations of applying image-level CLIP-based OOD detection methods to object-level scenarios. Building upon these insights, we propose RUNA, a novel framework that leverages a dual encoder architecture to capture rich contextual information and employs a regional uncertainty alignment mechanism to distinguish ID from OOD objects effectively. We introduce a few-shot fine-tuning approach that aligns region-level semantic representations to further improve the model's capability to discriminate between similar objects. Our experiments show that RUNA substantially surpasses state-of-the-art methods in object-level OOD detection, particularly in challenging scenarios with diverse and complex object instances. Jinggang Chen, Xiaoyang Qu, Guokuan Li, Kai Lu 0002, Jiguang Wan 0001, Jing Xiao 0006, Jianzong Wang |
AAAI | 3 |
| 2025 | Hierarchical-Task-Aware Multi-modal Mixture of Incremental LoRA Experts for Embodied Continual LearningabstractPrevious continual learning setups for embodied intelligence focused on executing low-level actions based on human commands, neglecting the ability to learn high-level planning and multi-level knowledge.To address these issues, we propose the Hierarchical Embodied Continual Learning Setups (HEC) that divide the agent's continual learning process into two layers: high-level instructions and low-level actions, and define five embodied continual learning sub-setups.Building on these setups, we introduce the Task-aware Mixture of Incremental LoRA Experts (Task-aware MoILE) method.This approach achieves task recognition by clustering visual-text embeddings and uses both a task-level router and a token-level router to select the appropriate LoRA experts.To effectively address the issue of catastrophic forgetting, we apply Singular Value Decomposition (SVD) to the LoRA parameters obtained from prior tasks, preserving key components while orthogonally training the remaining parts.The experimental results show that our method stands out in reducing the forgetting of old tasks compared to other methods, effectively supporting agents in retaining prior knowledge while continuously learning new tasks. Ziqi Jia, An-Min Wang, Xiaoyang Qu, Xiaowen Yang, Jianzong Wang |
ACL (1) | 3 |
| 2025 | MoQAE: Mixed-Precision Quantization for Long-Context LLM Inference via Mixture of Quantization-Aware ExpertsabstractOne of the primary challenges in optimizing large language models (LLMs) for long-context inference lies in the high memory consumption of the Key-Value (KV) cache. Existing approaches, such as quantization, have demonstrated promising results in reducing memory usage. However, current quantization methods cannot take both effectiveness and efficiency into account. In this paper, we propose MoQAE, a novel mixed-precision quantization method via mixture of quantization-aware experts. First, we view different quantization bit-width configurations as experts and use the traditional mixture of experts (MoE) method to select the optimal configuration. To avoid the inefficiency caused by inputting tokens one by one into the router in the traditional MoE method, we input the tokens into the router chunk by chunk. Second, we design a lightweight router-only fine-tuning process to train MoQAE with a comprehensive loss to learn the trade-off between model accuracy and memory usage. Finally, we introduce a routing freezing (RF) and a routing sharing (RS) mechanism to further reduce the inference overhead. Extensive experiments on multiple benchmark datasets demonstrate that our method outperforms state-of-the-art KV cache quantization approaches in both efficiency and effectiveness. Haocheng Lu, Xiaoyang Qu, Kai Lu 0002, Jiguang Wan 0001, Jianzong Wang |
ACL (1) | 3 |
| 2025 | Publicly Verifiable Private Information Retrieval Protocols Based on Function Secret Sharing
Lingwei Kong, Xiaoyang Qu, Jianzong Wang |
Inscrypt (3) | 4 |
| 2025 | Cocktail: Chunk-Adaptive Mixed-Precision Quantization for Long-Context LLM InferenceabstractRecently, large language models (LLMs) have been able to handle longer and longer contexts. However, a context that is too long may cause intolerant inference latency and GPU memory usage. Existing methods propose mixed-precision quantization to the key-value (KV) cache in LLMs based on token granularity, which is time-consuming in the search process and hardware inefficient during computation. This paper introduces a novel approach called Cocktail, which employs chunk-adaptive mixed-precision quantization to optimize the KV cache. Cocktail consists of two modules: chunk-level quantization search and chunk-level KV cache computation. Chunk-level quantization search determines the optimal bitwidth configuration of the KV cache chunks quickly based on the similarity scores between the corresponding context chunks and the query, maintaining the model accuracy. Furthermore, chunk-level KV cache computation reorders the KV cache chunks before quantization, avoiding the hardware inefficiency caused by mixed-precision quantization in inference computation. Extensive experiments demonstrate that Cocktail outperforms state-of-the-art KV cache quantization methods on various models and datasets. Our code is presented on https://github.com/Sullivan12138/Cocktail. Xiaoyang Qu, Jiguang Wan 0001, Jianzong Wang |
DATE | 3 |
| 2025 | Graph Contrastive Learning with Decoupled AugmentationabstractGraph contrastive learning based on augmentation strategies has recently demonstrated remarkable performance. Existing methods typically jointly leverage attribute and structural augmentations to generate graph views, learning data invariance information through contrasting sample pairs. However, this joint approach may deviate from the expectation of semantically similar before and after augmentation. The propagation of attribute information in graphs usually occurs through their structure, meaning that structural and attribute augmentations can interfere with each other and potentially distort the graph’s semantics. To address this, we propose a decoupled augmentation framework for graph contrastive learning, which eliminates the mutual interference between the two levels of augmentation while fully exploring graph information. Specifically, our framework employs separate encoders to learn data invariance under different augmentation levels, and it considers the positive gains generated between these levels. Experimental results on five public datasets show that the proposed method is more competitive than state-of-the-art approaches. Shihao Gao, Caoshuo Li, Cunli Mao, Xulong Zhang 0001, Xiaoyang Qu, Taisong Jin, Jianzong Wang |
ICASSP | 5 |
| 2025 | CycleFlow: Leveraging Cycle Consistency in Flow Matching for Speaker Style AdaptationabstractVoice Conversion (VC) aims to convert the style of a source speaker, such as timbre and pitch, to the style of any target speaker while preserving the linguistic content. However, the ground truth of the converted speech does not exist in a non-parallel VC scenario, which induces the train-inference mismatch problem. Moreover, existing methods still have an inaccurate pitch and low speaker adaptation quality, there is a significant disparity in pitch between the source and target speaker style domains. As a result, the models tend to generate speech with hoarseness, posing challenges in achieving high-quality voice conversion. In this study, we propose CycleFlow, a novel VC approach that leverages cycle consistency in conditional flow matching (CFM) for speaker timbre adaptation training on non-parallel data. Furthermore, we design a Dual-CFM based on VoiceCFM and PitchCFM to generate speech and improve speaker pitch adaptation quality. Experiments show that our method can significantly improve speaker similarity, generating natural and higher-quality speech. Ziqi Liang, Xulong Zhang 0001, Xiaoyang Qu, Weifeng Zhao, Jianzong Wang |
ICASSP | 4 |
| 2025 | PointActionCLIP: Preventing Transfer Degradation in Point Cloud Action Recognition with a Triple-Path CLIPabstractDirectly applying CLIP to point cloud action recognition can cause severe accuracy collapse. In this paper, we propose PointActionCLIP, which successfully prevents this transfer degradation with a triplepath CLIP, including the image path, the sequence path, and the label path. Specifically, the image path projects the 3D point cloud sequence onto a 2D image sequence and uses a visual encoder to extract its feature. It also captures the temporal feature of the image sequence with a temporal encoding transformer. The sequence path adopts a pretrained sequence encoder to encode the original point cloud sequence to obtain its spatiotemporal feature. The label path encodes the candidate labels with a text encoder. Finally, we fuse the output of the three paths to obtain the predicted action label. Extensive experiments validate that PointActionCLIP outperforms state-of-the-art (SOTA) methods. Shenglin He, Xiaoyang Qu, Jiguang Wan 0001, Jianzong Wang |
ICASSP | 3 |
| 2025 | VisTa: Visual-contextual and Text-augmented Zero-shot Object-level OOD DetectionabstractAs object detectors are increasingly deployed as black-box cloud services or pre-trained models with restricted access to the original training data, the challenge of zero-shot object-level out-of-distribution (OOD) detection arises. This task becomes crucial in ensuring the reliability of detectors in open-world settings. While existing methods have demonstrated success in image-level OOD detection using pre-trained vision-language models like CLIP, directly applying such models to object-level OOD detection presents challenges due to the loss of contextual information and reliance on image-level alignment. To tackle these challenges, we introduce a new method that leverages visual prompts and text-augmented in-distribution (ID) space construction to adapt CLIP for zero-shot object-level OOD detection. Our method preserves critical contextual information and improves the ability to differentiate between ID and OOD objects, achieving competitive performance across different benchmarks. Xiaoyang Qu, Guokuan Li, Jiguang Wan 0001, Jianzong Wang |
ICASSP | 2 |
| 2025 | Federated Domain Generalization with Domain-Specific Soft Prompts GenerationabstractPrompt learning has become an efficient paradigm for adapting CLIP to downstream tasks. Compared with traditional fine-tuning, prompt learning optimizes a few parameters yet yields highly competitive results, especially appealing in federated learning for computational efficiency. engendering domain shift among clients and posing a formidable challenge for downstream-task adaptation. Existing federated domain generalization (FDG) methods based on prompt learning typically learn soft prompts from training samples, replacing manually designed prompts to enhance the generalization ability of federated models. However, these learned prompts exhibit limited diversity and tend to ignore information from unknown domains. We propose a novel and effective method from a generative perspective for handling FDG tasks, namely federated domain generalization with domain-specific soft prompts generation (FedDSPG). Specifically, during training, we introduce domain-specific soft prompts (DSPs) for each domain and integrate content and domain knowledge into the generative model among clients. In the inference phase, the generator is utilized to obtain DSPs for unseen target domains, thus guiding downstream tasks in unknown domains. Comprehensive evaluations across several public datasets confirm that our method outperforms existing strong baselines in FDG, achieving state-of-the-art results. Jianhan Wu 0001, Xiaoyang Qu, Zhangcheng Huang 0002, Jianzong Wang |
ICCV | 2 |
| 2025 | MADLLM: Multivariate Anomaly Detection via Pre-trained LLMsabstractWhen applying pre-trained large language models (LLMs) to address anomaly detection tasks, the multivariate time series (MTS) modality of anomaly detection does not align with the text modality of LLMs. Existing methods simply transform the MTS data into multiple univariate time series sequences, which can cause many problems. This paper introduces MADLLM, a novel multivariate anomaly detection method via pre-trained LLMs. We design a new triple encoding technique to align the MTS modality with the text modality of LLMs. Specifically, this technique integrates the traditional patch embedding method with two novel embedding approaches: (i) Skip Embedding, which alters the order of patch processing in traditional methods to help LLMs retain knowledge of previous features, and (ii) Feature Embedding, which leverages contrastive learning to allow the model to better understand the correlations between different features. Experimental results demonstrate that our method outperforms state-of-the-art methods in various public anomaly detection datasets. Xiaoyang Qu, Kai Lu 0002, Jiguang Wan 0001, Guokuan Li, Jianzong Wang |
ICME | 2 |
| 2025 | Generalized Audio Deepfake Detection Using Frame-level Latent Information EntropyabstractGeneralizability, the capacity of a robust model to perform effectively on unseen data, is crucial for audio deepfake detection due to the rapid evolution of text-to-speech (TTS) and voice conversion (VC) technologies. A promising approach to differentiate between bonafide and spoof samples lies in identifying intrinsic disparities to enhance model generalizability. From an information-theoretic perspective, we hypothesize the information content is one of the intrinsic differences: bonafide sample represents a dense, information-rich sampling of the real world, whereas spoof sample is typically derived from lower-dimensional, less informative representations. To implement this, we introduce frame-level latent information entropy detector(f-InfoED), a framework that extracts distinctive information entropy from latent representations at the frame level to identify audio deepfakes. Furthermore, we present AdaLAM, which extends large pre-trained audio models with trainable adapters for enhanced feature extraction. To facilitate comprehensive evaluation, the audio deepfake forensics 2024 (ADFF 2024) dataset was built by the latest TTS and VC methods. Extensive experiments demonstrate that our proposed approach achieves state-of-the-art performance and exhibits remarkable generalization capabilities. Further analytical studies confirms the efficacy of AdaLAM in extracting discriminative audio features and f-InfoED in leveraging latent entropy information for more generalized deepfake detection. Botao Zhao 0001, Zuheng Kang, Yayun He, Xiaoyang Qu, Junqing Peng, Jing Xiao 0006, Jianzong Wang |
ICME | 4 |
| 2025 | Turbo-TTS: Enhancing Diffusion Model TTS with an Improved ODE Solver
Xulong Zhang 0001, Xiaoyang Qu, Hui Tian 0002, Jianzong Wang |
ICONIP (1) | 3 |
| 2025 | Enhancing Multi-Agent Systems via Reinforcement Learning with LLM-Based Planner and Graph-Based PolicyabstractMulti-agent systems (MAS) have shown great potential in executing complex tasks, but coordination and safety remain significant challenges. Multi-Agent Reinforcement Learning (MARL) offers a promising framework for agent collaboration, but it faces difficulties in handling complex tasks and designing reward functions. The introduction of Large Language Models (LLMs) has brought stronger reasoning and cognitive abilities to MAS, but existing LLM-based systems struggle to respond quickly and accurately in dynamic environments. To address these challenges, we propose LLM-based Graph Collaboration MARL (LGC-MARL), a framework that efficiently combines LLMs and MARL. This framework decomposes complex tasks into executable subtasks and achieves efficient collaboration among multiple agents through graph-based coordination. Specifically, LGC-MARL consists of two main components: an LLM planner and a graph-based collaboration meta policy. The LLM planner transforms complex task instructions into a series of executable subtasks, evaluates the rationality of these subtasks using a critic model, and generates an action dependency graph. The graph-based collaboration meta policy facilitates communication and collaboration among agents based on the action dependency graph, and adapts to new task environments through meta-learning. Experimental results on the AI2-THOR simulation platform demonstrate the superior performance and scalability of LGC-MARL in completing various complex tasks. Ziqi Jia, Xiaoyang Qu, Jianzong Wang |
ICRA | 3 |
| 2025 | BAGNet: A Boundary-Aware Graph Attention Network for 3D Point Cloud Semantic SegmentationabstractSince the point cloud data is inherently irregular and unstructured, point cloud semantic segmentation has always been a challenging task. The graph-based method attempts to model the irregular point cloud by representing it as a graph; however, this approach incurs substantial computational cost due to the necessity of constructing a graph for every point within a large-scale point cloud. In this paper, we observe that boundary points possess more intricate spatial structural information and develop a novel graph attention network known as the Boundary-Aware Graph attention Network (BAGNet). On one hand, BAGNet contains a boundary-aware graph attention layer (BAGLayer), which employs edge vertex fusion and attention coefficients to capture features of boundary points, reducing the computation time. On the other hand, BAGNet employs a lightweight attention pooling layer to extract the global feature of the point cloud to maintain model accuracy. Extensive experiments on standard datasets demonstrate that BAGNet outperforms state-of-the-art methods in point cloud semantic segmentation with higher accuracy and less inference time. Xiaoyang Qu, Kai Lu 0002, Jiguang Wan 0001, Shenglin He, Jianzong Wang |
IJCNN | 2 |
| 2025 | Rano: Restorable Speaker Anonymization via Conditional Invertible Neural NetworkabstractSpeech contains ample information, including the primary semantic content and information about the speaker, such as gender, age and health status. Speaker-dependent information partially carries personal privacy and has raised concerns about the protection of voice privacy. Speaker anonymization aims to conceal the speaker’s identity in speech, while preserving speaker-independent information to the greatest extent possible, and has become an increasingly important task in the field of speech. Existing research generally treats speaker anonymization as a downstream task of voice conversion, often employing speech representation disentanglement-based methods to separate speaker-dependent and speaker-independent information. However, speech representation disentanglement, especially for speaker-independent information, faces challenges such as information leakage or excessive disentangling, resulting in quality degradation. In this paper, we propose a speaker anonymization model called Rano, which does not rely on precise disentanglement. Rano employs a generative invertible neural network to forge anonymous speaker identities from keys and then uses the speaker embeddings as conditions to guide the speaker anonymization process via a conditional invertible neural network. Moreover, when the key is provided, lossless restoration from anonymized speech to original speech can be achieved via the reverse network, thereby expanding the application scenarios of Rano. Experiments demonstrate that the proposed model achieves comparable performance to existing state-of-the-art models. We also verify the security guaranteed by the key in the restoration process. Jianzong Wang, Xulong Zhang 0001, Xiaoyang Qu |
IJCNN | 3 |
| 2025 | Bridging the Modality Gap: Semantic-Calibrated Zero-shot Speech Emotion CaptioningabstractSpeech Emotion Captioning (SEC) has emerged as an increasingly prominent research area. The emotional content expressed through human speech is often intricate, making it difficult to fully capture with fixed categorical labels. Instead, describing these emotions using natural language can offer a more comprehensive representation. However, obtaining well-matched speech-caption datasets is challenging in practical scenarios, and existing SEC techniques often generate hallucinated content or miss fine-grained details when working with cross-domain, unpaired data. To address these challenges, we introduce SeCCap, a Semantic-Calibrated Zero-shot Speech Emotion Captioning framework built upon large language models (LLMs). SeCCap exhibits three key features: 1) Zero-shot inference: It generates speech emotion captions without requiring training on paired speech-caption datasets. 2) Bridging the modality gap: It employs caption-only training and semantic-calibrated cross-modal mapping to enhance fine-grained content and reduce factual hallucinations during zero-shot SEC. Experimental results show that SeCCap outperforms other state-of-the-art models in zero-shot SEC tasks. Jianzong Wang, Xulong Zhang 0001, Xiaoyang Qu |
IJCNN | 3 |
| 2025 | Data-free Black-box Knowledge AmalgamationabstractA massive number of well-trained models with promising performances have been released nowadays; exploring reusing them would benefit the community. Some recent works propose to amalgamate multiple models’ pre-learned knowledge and transfer into a single model. They distill knowledge from teachers’ intermediate layers and use unlabeled data as the distilling source. However, in many real-world cases, teachers’ model architectures vary. Thus, aligning the intermediate layers takes a lot of work. Also, the unlabeled data is required to belong to the same domain where teachers pre-learned, which likely is confidential and unavailable. Both constraints would limit the practicability of knowledge amalgamation. To tackle these problems, we propose to treat teacher models as black boxes and only amalgamate teachers’ responses in a data-free manner, thus relaxing both constraints. Further, we validate and address the unfairness and uncertainty issue in the amalgamated response. By entropy transform and subjective re-weighting, we build a confident amalgamated response that can better guide knowledge transfer. Extensive experiments demonstrate that our method can significantly improve performance in various heterogeneous settings. Jianzong Wang, Chendong Zhao, Xiaoyang Qu |
IJCNN | 4 |
| 2025 | Knowledge distillation for financial large language models: a systematic review of strategies, applications, and evaluationabstractFinancial large language models (FinLLMs) offer immense potential for financial applications. While excessive deployment expenditures and considerable inference latency constitute major obstacles, as a prominent compression methodology, knowledge distillation (KD) offers an effective solution to these difficulties. A comprehensive survey is conducted in this work on how KD interacts with FinLLMs, covering three core aspects: strategy, application, and evaluation. At the strategy level, this review introduces a structured taxonomy to comparatively analyze existing distillation pathways. At the application level, this review puts forward a logical upstream–midstream–downstream framework to systematically explain the practical value of distilled models in the financial field. At the evaluation level, to tackle the absence of standards in the financial field, this review constructs a comprehensive evaluation framework that proceeds from multiple dimensions such as financial accuracy, reasoning fidelity, and robustness. In summary, this research aims to provide a clear roadmap for this interdisciplinary field, to accelerate the development of distilled FinLLMs. Xulong Zhang 0001, Xiaoyang Qu, Junfei Xie, Jianzong Wang |
Frontiers Inf. Technol. Electron. Eng. | 3 |
| 2024 | Beyond Aggregation: Efficient Federated Model Consolidation with Heterogeneity-Adaptive Weights DiffusionabstractAs the Internet of Things (IoT) evolves, the need for enhanced data-sharing to improve edge device performance has led to the adoption of Federated Learning (FL) for data privacy and optimized data utilization. However, communication costs in FL remain a significant challenge. Traditional methods focus on client enhancements but overlook server-side aggregation, potentially increasing client computation loads. In response, we introduce a novel method, FedDiff, which utilizes diffusion models for generating model weights on FL servers, replacing traditional aggregation methods. Our approach, tailored for heterogeneous environments, significantly improves communication efficiency, achieving faster convergence and robust performance against weight noise in rigorous tests. Jiaqi Li 0028, Xiaoyang Qu, Wenbo Ding 0001, Zihao Zhao 0001, Jianzong Wang |
CIKM | 2 |
| 2024 | Value-Driven Mixed-Precision Quantization for Patch-Based Inference on MicrocontrollersabstractDeploying neural networks on microcontroller units (MCUs) presents substantial challenges due to their constrained computation and memory resources. Previous researches have explored patch-based inference as a strategy to conserve memory without sacrificing model accuracy. However, this technique suffers from severe redundant computation overhead, leading to a substantial increase in execution latency. A feasible solution to address this issue is mixed-precision quantization, but it faces the challenges of accuracy degradation and a time-consuming search time. In this paper, we propose QuantMCU, a novel patch-based inference method that utilizes value-driven mixed-precision quantization to reduce redundant computation. We first utilize value-driven patch classification (VDPC) to maintain the model accuracy. VDPC classifies patches into two classes based on whether they contain outlier values. For patches containing outlier values, we apply 8-bit quantization to the feature maps on the dataflow branches that follow. In addition, for patches without outlier values, we utilize value-driven quantization search (VDQS) on the feature maps of their following dataflow branches to reduce search time. Specifically, VDQS introduces a novel quantization search metric that takes into account both computation and accuracy, and it employs entropy as an accuracy representation to avoid additional training. VDQS also adopts an iterative approach to determine the bitwidth of each feature map to further accelerate the search process. Experimental results on real-world MCU devices show that QuantMCU can reduce computation by 2.2x on average while maintaining comparable model accuracy compared to the state-of-the-art patch-based inference methods. Shenglin He, Kai Lu 0002, Xiaoyang Qu, Guokuan Li, Jiguang Wan 0001, Jianzong Wang, Jing Xiao 0006 |
DATE | 4 |
| 2024 | Incremental Label Distribution Learning with Scalable Graph Convolutional NetworksabstractLabel Distribution Learning (LDL) is an effective approach for handling label ambiguity, as it can analyze all labels at once and indicate the extent to which each label describes a given sample. Most existing LDL methods consider the number of labels to be static. However, in various LDL-specific contexts (e.g., disease diagnosis), the label count grows over time (such as the discovery of new diseases), a factor that existing methods overlook. Learning samples with new labels directly means learning all labels at once, thus wasting more time on the old labels and even risking overfitting the old labels. At the same time, learning new labels by the LDL model means reconstructing the inter-label relationships. How to make use of constructed relationships is also a crucial challenge. To tackle these challenges, we introduce Incremental Label Distribution Learning (ILDL), analyze its key issues regarding training samples and inter-label relationships, and propose Scalable Graph Label Distribution Learning (SGLDL) as a practical framework for implementing ILDL. Specifically, in SGLDL, we develop a New-label-aware Gradient Compensation Loss to speed up the learning of new labels and represent inter-label relationships as a graph to reduce the time required to reconstruct inter-label relationships. Experimental results on the classical LDL dataset show the clear advantages of unique algorithms and illustrate the importance of a dedicated design for the ILDL problem. Ziqi Jia, Xiaoyang Qu, Jianzong Wang |
HPCC | 2 |
| 2024 | INCPrompt: Task-Aware Incremental Prompting for Rehearsal-Free Class-Incremental LearningabstractThis paper introduces INCPrompt, an innovative continual learning solution that effectively addresses catastrophic forgetting. INCPrompt’s key innovation lies in its use of adaptive key-learner and task-aware prompts that capture task-relevant information. This unique combination encapsulates general knowledge across tasks and encodes task-specific knowledge. Our comprehensive evaluation across multiple continual learning benchmarks demonstrates INCPrompt’s superiority over existing algorithms, showing its effectiveness in mitigating catastrophic forgetting while maintaining high performance. These results highlight the significant impact of task-aware incremental prompting on continual learning performance. Xiaoyang Qu, Jing Xiao 0006, Bokui Chen, Jianzong Wang |
ICASSP | 2 |
| 2024 | P2DT: Mitigating Forgetting in Task-Incremental Learning with Progressive Prompt Decision TransformerabstractCatastrophic forgetting poses a substantial challenge for managing intelligent agents controlled by a large model, causing performance degradation when these agents face new tasks. In our work, we propose a novel solution - the Progressive Prompt Decision Transformer (P2DT). This method enhances a transformer-based model by dynamically appending decision tokens during new task training, thus fostering task-specific policies. Our approach mitigates forgetting in continual and offline reinforcement learning scenarios. Moreover, P2DT leverages trajectories collected via traditional reinforcement learning from all tasks and generates new taskspecific tokens during training, thereby retaining knowledge from previous studies. Preliminary results demonstrate that our model effectively alleviates catastrophic forgetting and scales well with increasing task environments. Xiaoyang Qu, Jing Xiao 0006, Bokui Chen, Jianzong Wang |
ICASSP | 2 |
| 2024 | Enhancing Anomalous Sound Detection with Multi-Level Memory BankabstractAbnormal sound detection (ASD) is crucial for the timely detection of machine faults in industrial scenarios and has emerged as a popular topic. However, exhaustively collecting all ever-changing anomalous samples is impractical for the associated time and cost. Under unsupervised conditions, identifying rare or even unseen abnormal sounds from a large set of normal samples is a notable challenge in the real-world setting. To address this, we propose a novel ASD method based on a multi-level memory bank to estimate the distribution of normal samples in the latent space. We employ a distance-based metric to distinguish inliers from outliers, leveraging high, mid, and low-level features to improve accuracy. We also propose an acoustic-aware farthest embedding sampling algorithm for inference acceleration and memory bank reduction. Experimental results demonstrate our method outperforms existing methods for anomaly detection. Additionally, we analyze the effect of multilevel and acoustic-aware farthest embedding sampling methods, respectively. Baoping Deng, Jinggang Chen, Zhenhou Hong, Xiaoyang Qu, Guokuan Li, Jiguang Wan 0001, Jianzong Wang |
IJCNN | 4 |
| 2024 | PRENet: A Plane-Fit Redundancy Encoding Point Cloud Sequence Network for Real-Time 3D Action RecognitionabstractRecognizing human actions from point cloud sequence has attracted tremendous attention from both academia and industry due to its wide applications. However, most previous studies on point cloud action recognition typically require complex networks to extract intra-frame spatial features and inter-frame temporal features, resulting in an excessive number of redundant computations. This leads to high latency, rendering them impractical for real-world applications. To address this problem, we propose a Plane-Fit Redundancy Encoding point cloud sequence network named PRENet. The primary concept of our approach involves the utilization of plane fitting to mitigate spatial redundancy within the sequence, concurrently encoding the temporal redundancy of the entire sequence to minimize redundant computations. Specifically, our network comprises two principal modules: a Plane-Fit Embedding module and a Spatio-Temporal Consistency Encoding module. The Plane-Fit Embedding module capitalizes on the observation that successive point cloud frames exhibit unique geometric features in physical space, allowing for the reuse of spatially encoded data for temporal stream encoding. The Spatio-Temporal Consistency Encoding module amalgamates the temporal structure of the temporally redundant part with its corresponding spatial arrangement, thereby enhancing recognition accuracy. We have done numerous experiments to verify the effectiveness of our network. The experimental results demonstrate that our method achieves almost identical recognition accuracy while being nearly four times faster than other state-of-the-art methods. Shenglin He, Xiaoyang Qu, Jiguang Wan 0001, Guokuan Li, Jianzong Wang |
IJCNN | 2 |
| 2024 | Task-agnostic Decision Transformer for Multi-type Agent Control with Federated Split TrainingabstractWith the rapid advancements in artificial intelligence, the development of knowledgeable and personalized agents has become increasingly prevalent. However, the inherent variability in state variables and action spaces among personalized agents poses significant aggregation challenges for traditional federated learning algorithms. To tackle these challenges, we introduce the Federated Split Decision Transformer (FSDT), an innovative framework designed explicitly for AI agent decision tasks. The FSDT framework excels at navigating the intricacies of personalized agents by harnessing distributed data for training while preserving data privacy. It employs a two-stage training process, with local embedding and prediction models on client agents and a global transformer decoder model on the server. Our comprehensive evaluation using the benchmark D4RL dataset highlights the superior performance of our algorithm in federated split learning for personalized agents, coupled with significant reductions in communication and computational overhead compared to traditional centralized training approaches. The FSDT framework demonstrates strong potential for enabling efficient and privacy-preserving collaborative learning in applications such as autonomous driving decision systems. Our findings underscore the efficacy of the FSDT framework in effectively leveraging distributed offline reinforcement learning data to enable powerful multi-type agent decision systems. Bokui Chen, Xiaoyang Qu, Zhenhou Hong, Jing Xiao 0006, Jianzong Wang |
IJCNN | 3 |
| 2024 | Gecko: Resource-Efficient and Accurate Queries in Real-Time Video Streams at the EdgeabstractSurveillance cameras are ubiquitous nowadays and users’ increasing needs for accessing real-world information (e.g., finding abandoned luggage) have urged object queries in real-time videos. While recent real-time video query processing systems exhibit excellent performance, they lack utility in deployment in practice as they overlook some crucial aspects, including multi-camera exploration, resource contention, and content awareness. Motivated by these issues, we propose a framework Gecko, to provide resource-efficient and accurate real-time object queries of massive videos on edge devices. Gecko (i) obtains optimal models from the model zoo and assigns them to edge devices for executing current queries, (ii) optimizes resource usage of the edge cluster at runtime by dynamically adjusting the frame query interval of each video stream and forking/joining running models on edge devices, and (iii) improves accuracy in changing video scenes by fine-grained stream transfer and continuous learning of models. Our evaluation with real-world video streams and queries shows that Gecko achieves up to 2x more resource efficiency gains and increases overall query accuracy by at least 12% compared with prior work, further delivering excellent scalability for practical deployment. Liang Wang 0057, Xiaoyang Qu, Jianzong Wang, Guokuan Li, Jiguang Wan 0001, Song Guo 0001, Jing Xiao 0006 |
INFOCOM | 2 |
| 2024 | Retrieval-Augmented Audio Deepfake DetectionabstractWith recent advances in speech synthesis including text-to-speech (TTS) and voice conversion (VC) systems enabling the generation of ultra-realistic audio deepfakes, there is growing concern about their potential misuse. However, most deepfake (DF) detection methods rely solely on the fuzzy knowledge learned by a single model, resulting in performance bottlenecks and transparency issues. Inspired by retrieval-augmented generation (RAG), we propose a retrieval-augmented detection (RAD) framework that augments test samples with similar retrieved samples for enhanced detection. We also extend the multi-fusion attentive classifier to integrate it with our proposed RAD framework. Extensive experiments show the superior performance of the proposed RAD framework over baseline methods, achieving state-of-the-art results on the ASVspoof 2021 DF set and competitive results on the 2019 and 2021 LA sets. Further sample analysis indicates that the retriever consistently retrieves samples mostly from the same speaker with acoustic characteristics highly consistent with the query audio, thereby improving detection performance. Zuheng Kang, Yayun He, Botao Zhao 0001, Xiaoyang Qu, Junqing Peng, Jing Xiao 0006, Jianzong Wang |
ICMR | 4 |
| 2024 | FormerReckoning: Physics Inspired Transformer for Accurate Inertial NavigationabstractAlthough modern localization methods have achieved remarkable accuracy with various sensors, there are still some circumstances where only proprioceptive sensing works (Inertial Navigation). However, localization and navigation using only IMU sensors (costing less than $1000) still face significant challenges such as low accuracy and large cumulative errors when using traditional filter methods. Furthermore, AI-based approaches, while promising, often yield unpredictable and unreliable outputs. This paper proposes FormerReckoning, an inertial localization estimation framework for wheeled robotics that incorporates physical prompts into a Transformer framework to enhance translation estimation accuracy. Our tests show that FormerReckoning not only reduces mean translation errors to 0.72% but also surpasses all baseline models in performance, demonstrating its potential to provide reliable and precise localization in a cost-effective manner. Jiaqi Li 0028, Chenyu Zhao 0002, Yuzhu Mao, Xinlei Chen, Wenbo Ding 0001, Xiaoyang Qu, Jianzong Wang |
MobiCom | 6 |
| 2024 | Rule-Guided Counterfactual Explainable RecommendationabstractTo empower the trust of current recommender systems, the counterfactual explanation (CE) method is adopted to generate the counterfactual instance for each input and take their changes causing the different outcomes as the explanation. Although promising results have been achieved by existing CE-based methods, we propose to generate the attribute-oriented counterfactual explanation. Different from them, we aim to generate the counterfactual instance by performing the intervention on the attributes, and then build an attribute-oriented counterfactual explainable recommender system. Considering the correlation and categorical values of attributes, how to efficiently generate the reliable counterfactual instances on the attributes challenges us. To alleviate such a problem, we propose to extract the decision rules over the attributes to guide the attribute-oriented counterfactual generation. Specifically, we adopt the gradient boosting decision tree (GBDT) to pre-build the decision rules over the attributes and develop a Rule-guided Counterfactual Explainable Recommendation model (RCER) to predict the user-item interaction and generate the counterfactual instances for the user-item pairs. We finally conduct extensive experiments on four publicly datasets, including NYC, LON, Amazon, and Movielens datasets. Experimental results have qualitatively and quantitatively justified the superiority of our model over existing cutting-edge baselines. We release the code:https://github.com/quxiaoyang0zero/RCER. Yinwei Wei, Xiaoyang Qu, Xiang Wang 0010, Yunshan Ma 0002, Liqiang Nie, Tat-Seng Chua |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2023 | Shoggoth: Towards Efficient Edge-Cloud Collaborative Real-Time Video Inference via Adaptive Online LearningabstractThis paper proposes Shoggoth, an efficient edge-cloud collaborative architecture, for boosting inference performance on real-time video of changing scenes. Shoggoth uses online knowledge distillation to improve the accuracy of models suffering from data drift and offloads the labeling process to the cloud, alleviating constrained resources of edge devices. At the edge, we design adaptive training using small batches to adapt models under limited computing power, and adaptive sampling of training frames for robustness and reducing bandwidth. The evaluations on the realistic dataset show 15%–20% model accuracy improvement compared to the edge-only strategy and fewer network costs than the cloud-only strategy. Liang Wang 0057, Kai Lu 0002, Xiaoyang Qu, Jianzong Wang, Jiguang Wan 0001, Guokuan Li, Jing Xiao 0006 |
DAC | 4 |
| 2023 | Detecting Out-of-Distribution Examples Via Class-Conditional Impressions ReappearingabstractOut-of-distribution (OOD) detection aims at enhancing standard deep neural networks to distinguish anomalous inputs from original training data. Previous progress has introduced various approaches where the in-distribution training data and even several OOD examples are prerequisites. However, due to privacy and security, auxiliary data tends to be impractical in a real-world scenario. In this paper, we propose a data-free method without training on natural data, called Class-Conditional Impressions Reappearing (C2IR), which utilizes image impressions from the fixed model to recover class-conditional feature statistics. Based on that, we introduce Integral Probability Metrics to estimate layer-wise class-conditional deviations and obtain layer weights by Measuring Gradient-based Importance (MGI). The experiments verify the effectiveness of our method and indicate that C2IR outperforms other post-hoc methods and reaches comparable performance to the full access (ID and OOD) detection method, especially in the far-OOD dataset (SVHN). Jinggang Chen, Xiaoyang Qu, Jianzong Wang, Jiguang Wan 0001, Jing Xiao 0006 |
ICASSP | 2 |
| 2023 | Feature-Rich Audio Model Inversion for Data-Free Knowledge Distillation Towards General Sound ClassificationabstractData-Free Knowledge Distillation (DFKD) has recently attracted growing attention in the academic community, especially with major breakthroughs in computer vision. Despite promising results, the technique has not been well applied to audio and signal processing. Due to the variable duration of audio signals, it has its own unique way of modeling. In this work, we propose feature-rich audio model inversion (FRAMI), a data-free knowledge distillation framework for general sound classification tasks. It first generates high-quality and feature-rich Mel-spectrograms through a feature-invariant contrastive loss. Then, the hidden states before and after the statistics pooling layer are reused when knowledge distillation is performed on these feature-rich samples. Experimental results on the Urbansound8k, ESC-50, and audioMNIST datasets demonstrate that FRAMI can generate feature-rich samples. Meanwhile, the accuracy of the student model is further improved by reusing the hidden state and significantly outperforms the baseline method. Zuheng Kang, Yayun He, Jianzong Wang, Junqing Peng, Xiaoyang Qu, Jing Xiao 0006 |
ICASSP | 5 |
| 2023 | EdgeMA: Model Adaptation System for Real-Time Video Analytics on Edge Devices
Liang Wang 0057, Xiaoyang Qu, Jianzong Wang, Jiguang Wan 0001, Guokuan Li, Kaiyu Hu, Guilin Jiang, Jing Xiao 0006 |
ICONIP (1) | 3 |
| 2023 | FedET: A Communication-Efficient Federated Class-Incremental Learning Framework Based on Enhanced TransformerabstractFederated Learning (FL) has been widely concerned for it enables decentralized learning while ensuring data privacy. However, most existing methods unrealistically assume that the classes encountered by local clients are fixed over time. After learning new classes, this impractical assumption will make the model's catastrophic forgetting of old classes significantly severe. Moreover, due to the limitation of communication cost, it is challenging to use large-scale models in FL, which will affect the prediction accuracy. To address these challenges, we propose a novel framework, Federated Enhanced Transformer (FedET), which simultaneously achieves high accuracy and low communication cost. Specifically, FedET uses Enhancer, a tiny module, to absorb and communicate new knowledge, and applies pre-trained Transformers combined with different Enhancers to ensure high precision on various tasks. To address local forgetting caused by new classes of new tasks and global forgetting brought by non-i.i.d class imbalance across different local clients, we proposed an Enhancer distillation method to modify the imbalance between old and new knowledge and repair the non-i.i.d. problem. Experimental results demonstrate that FedET's average accuracy on a representative benchmark dataset is 14.1% higher than the state-of-the-art method, while FedET saves 90% of the communication cost compared to the previous method. Xiaoyang Qu, Jianzong Wang, Jing Xiao 0006 |
IJCAI | 2 |
| 2023 | GAIA: Delving into Gradient-based Attribution Abnormality for Out-of-distribution DetectionabstractDetecting out-of-distribution (OOD) examples is crucial to guarantee the reliability and safety of deep neural networks in real-world settings. In this paper, we offer an innovative perspective on quantifying the disparities between in-distribution (ID) and OOD data---analyzing the uncertainty that arises when models attempt to explain their predictive decisions. This perspective is motivated by our observation that gradient-based attribution methods encounter challenges in assigning feature importance to OOD data, thereby yielding divergent explanation patterns. Consequently, we investigate how attribution gradients lead to uncertain explanation outcomes and introduce two forms of abnormalities for OOD detection: the zero-deflation abnormality and the channel-wise average abnormality. We then propose GAIA, a simple and effective approach that incorporates Gradient Abnormality Inspection and Aggregation. The effectiveness of GAIA is validated on both commonly utilized (CIFAR) and large-scale (ImageNet-1k) benchmarks. Specifically, GAIA reduces the average FPR95 by 23.10% on CIFAR10 and by 45.41% on CIFAR100 compared to advanced post-hoc methods. Jinggang Chen, Xiaoyang Qu, Jianzong Wang, Jiguang Wan 0001, Jing Xiao 0006 |
NeurIPS | 3 |
| 2022 | Pose Guided Human Image Synthesis with Partially Decoupled GAN
Jianhan Wu 0001, Shijing Si, Jianzong Wang, Xiaoyang Qu, Jing Xiao 0006 |
ACML | 4 |
| 2022 | Adaptive Sparse and Monotonic Attention for Transformer-based Automatic Speech RecognitionabstractThe Transformer architecture model, based on self-attention and multi-head attention, has achieved remarkable success in offline end-to-end Automatic Speech Recognition (ASR). However, self-attention and multi-head attention cannot be easily applied for streaming or online ASR. For self-attention in Transformer ASR, the softmax normalization function-based attention mechanism makes it impossible to highlight important speech information. For multi-head attention in Transformer ASR, it is not easy to model monotonic alignments in different heads. To overcome these two limits, we integrate sparse attention and monotonic attention into Transformer-based ASR. The sparse mechanism introduces a learned sparsity scheme to enable each self-attention structure to fit the corresponding head better. The monotonic attention deploys regularization to prune redundant heads for the multi-head attention structure. The experiments show that our method can effectively improve the attention mechanism on widely used benchmarks of speech recognition. Chendong Zhao, Jianzong Wang, Xiaoyang Qu, Haoqian Wang, Jing Xiao 0006 |
DSAA | 4 |
| 2022 | r-G2P: Evaluating and Enhancing Robustness of Grapheme to Phoneme Conversion by Controlled Noise Introducing and Contextual Information IncorporationabstractGrapheme-to-phoneme (G2P) conversion is the process of converting the written form of words to their pronunciations. It has an important role for text-to-speech (TTS) synthesis and automatic speech recognition (ASR) systems. In this paper, we aim to evaluate and enhance the robustness of G2P models. We show that neural G2P models are extremely sensitive to orthographical variations in graphemes like spelling mistakes. To solve this problem, we propose three controlled noise introducing methods to synthesize noisy training data. Moreover, we incorporate the contextual information with the baseline and propose a robust training strategy to stabilize the training process. The experimental results demonstrate that our proposed robust G2P model (r-G2P) outperforms the baseline significantly (-2.73% WER on Dict-based benchmarks and -9.09% WER on Real-world sources). Chendong Zhao, Jianzong Wang, Xiaoyang Qu, Haoqian Wang, Jing Xiao 0006 |
ICASSP | 3 |
| 2022 | Boosting StarGANs for Voice Conversion with Contrastive Discriminator
Shijing Si, Jianzong Wang, Xulong Zhang 0001, Xiaoyang Qu, Ning Cheng 0001, Jing Xiao 0006 |
ICONIP (2) | 4 |
| 2022 | Blur the Linguistic Boundary: Interpreting Chinese Buddhist Sutra in English via Neural Machine TranslationabstractBuddhism is an influential religion with a long-standing history and profound philosophy. Nowadays, more and more people worldwide aspire to learn the essence of Buddhism, attaching importance to Buddhism dissemination. However, Buddhist scriptures written in classical Chinese are obscure to most people and machine translation applications. For instance, general Chinese-English neural machine translation (NMT) fails in this domain. In this paper, we proposed a novel approach to building a practical NMT model for Buddhist scriptures. The performance of our translation pipeline acquired highly promising results in ablation experiments under three criteria. Denghao Li, Yuqiao Zeng, Jianzong Wang, Lingwei Kong, Zhangcheng Huang 0002, Ning Cheng 0001, Xiaoyang Qu, Jing Xiao 0006 |
ICTAI | 7 |
| 2022 | QSpeech: Low-Qubit Quantum Speech Application ToolkitabstractQuantum devices with low qubits are common in the Noisy Intermediate-Scale Quantum (NISQ) era. However, Quantum Neural Network (QNN) running on low-qubit quantum devices would be difficult since it is based on Variational Quantum Circuit (VQC), which requires many qubits. Therefore, it is critical to make QNN with VQC run on low-qubit quantum devices. In this study, we propose a novel VQC called the low-qubit VQC. VQC requires numerous qubits based on the input dimension; however, the low-qubit VQC with linear transformation can liberate this condition. Thus, it allows the QNN to run on low-qubit quantum devices for speech applications. Furthermore, as compared to the VQC, our proposed low-qubit VQC can stabilize the training process more. Based on the low-qubit VQC, we implement QSpeech11Our implementation of QSpeech is publicly available at https://github.com/zhenhouhong/QSpeech, a library for quick prototyping of hybrid quantum-classical neural networks in the speech field. It has numerous quantum neural layers and QNN models for speech applications. Experiments on Speech Command Recognition and Text-to-Speech show that our proposed low-qubit VQC outperforms VQC and is more stable. Zhenhou Hong, Jianzong Wang, Xiaoyang Qu, Chendong Zhao, Jing Xiao 0006 |
IJCNN | 3 |
| 2022 | Leveraging Causal Inference for Explainable Automatic Program RepairabstractDeep learning models have made significant progress in automatic program repair. However, the black-box nature of these methods has restricted their practical applications. To address this challenge, this paper presents an interpretable approach for program repair based on sequence-to-sequence models with causal inference and our method is called CPR, short for causal program repair. Our CPR can generate explanations in the process of decision making, which consists of groups of causally related input-output tokens. Firstly, our method infers these relations by querying the model with inputs disturbed by data augmentation. Secondly, it generates a graph over tokens from the responses and solves a partitioning problem to select the most relevant components. The experiments on four programming languages (Java, C, Python, and JavaScript) show that CPR can generate causal graphs for reasonable interpretations and boost the performance of bug fixing in automatic program repair. Jianzong Wang, Shijing Si, Zhitao Zhu, Xiaoyang Qu, Zhenhou Hong, Jing Xiao 0006 |
IJCNN | 4 |
| 2022 | DT-SV: A Transformer-based Time-domain Approach for Speaker VerificationabstractSpeaker verification (SV) aims to determine whether the speaker's identity of a test utterance is the same as the reference speech. In the past few years, extracting speaker embeddings using deep neural networks for SV systems has gone mainstream. Recently, different attention mechanisms and Transformer networks have been explored widely in SV fields. However, utilizing the original Transformer in SV directly may have frame-level information waste on output features, which could lead to restrictions on capacity and discrimination of speaker embeddings. Therefore, we propose an approach to derive utterance-level speaker embeddings via a Transformer architecture that uses a novel loss function named diffluence loss to integrate the feature information of different Transformer layers. Therein, the diffluence loss aims to aggregate frame-level features into an utterance-level representation, and it could be integrated into the Transformer expediently. Besides, we also introduce a learnable mel-fbank energy feature extractor named time-domain feature extractor that computes the mel-fbank features more precisely and efficiently than the standard mel-fbank extractor. Combining Diffluence loss and Time-domain feature extractor, we propose a novel Transformer-based time-domain SV model (DT-SV) with faster training speed and higher accuracy. Experiments indicate that our proposed model can achieve better performance in comparison with other models. Jianzong Wang, Zhenhou Hong, Chendong Zhao, Xiaoyang Qu, Jing Xiao 0006 |
IJCNN | 5 |
| 2022 | Adaptive Few-Shot Learning Algorithm for Rare Sound Event DetectionabstractSound event detection is to infer the event by understanding the surrounding environmental sounds. Due to the scarcity of rare sound events, it becomes challenging for the well-trained detectors which have learned too much prior knowledge. Meanwhile, few-shot learning methods promise a good generalization abaility when facing a new limited-data task. Recent approaches have achieved promising results in this field. However, these approaches treat each support example independently, ignoring the information of other examples from the whole task. Because of this, most of previous methods are constrained to generate a same feature embedding for all test-time tasks, which is not adaptive to each inputted data. In this work, we propose a novel task-adaptive module which is easy to plant into any metric-based few-shot learning frameworks. The module could identify the task-relevant feature dimension. Incorporating our module improves the performance considerably on two datasets over baseline methods, especially for the transductive propagation network. Such as +6.8% for 5-way 1-shot accuracy on ESC-50, and +5.9% on noiseESC-50. We investigate our approach in the domain-mismatch setting and also achieve better results than previous methods. Chendong Zhao, Jianzong Wang, Leilai Li, Xiaoyang Qu, Jing Xiao 0006 |
IJCNN | 4 |
| 2022 | Learning Invariant Representation and Risk Minimized for Unsupervised Accent Domain AdaptationabstractUnsupervised representation learning for speech audios attained impressive performances for speech recognition tasks, particularly when annotated speech is limited. However, the unsupervised paradigm needs to be carefully designed and little is known about what properties these representations acquire. There is no guarantee that the model learns meaningful representations for valuable information for recognition. Moreover, the adaptation ability of the learned representations to other domains still needs to be estimated. In this work, we explore learning domain-invariant representations via a direct mapping of speech representations to their corresponding high-level linguistic informations. Results prove that the learned latents not only capture the articulatory feature of each phoneme but also enhance the adaptation ability, outperforming the baseline largely on accented benchmarks. Chendong Zhao, Jianzong Wang, Xiaoyang Qu, Haoqian Wang, Jing Xiao 0006 |
SLT | 3 |
| 2021 | Enhancing Data-Free Adversarial Distillation with Activation Regularization and Virtual InterpolationabstractKnowledge distillation refers to a technique of transferring the knowledge from a large learned model or an ensemble of learned models to a small model. This method relies on access to the original training set, which might not always be available. A possible solution is a data-free adversarial distillation framework, which deploys a generative network to transfer the teacher model’s knowledge to the student model. However, the data generation efficiency is low in the data-free adversarial distillation. We add an activation regularizer and a virtual interpolation method to improve the data generation efficiency. The activation regularizer enables the students to match the teacher’s predictions close to activation boundaries and decision boundaries. The virtual interpolation method can generate virtual samples and labels in-between decision boundaries. Our experiments show that our approach surpasses state-of-the-art data-free distillation methods. The student model can achieve 95.42% accuracy on CIFAR-10 and 77.05% accuracy on CIFAR-100 without any original training data. Our model’s accuracy is 13.8% higher than the state-of-the-art data-free method on CIFAR-100. Xiaoyang Qu, Jianzong Wang, Jing Xiao 0006 |
ICASSP | 1 |
| 2021 | Quantum Convolutional Neural Network on Protein Distance PredictionabstractProteins are linear polymers that fold into an incredible variety of three-dimensional structures that enable sophisticated functionality for biology. Predicting protein distance with high precision remains challenging, particularly for small protein families. As deep learning achieves remarkable success in many areas, deep learning also allows scientists to predict proteins' three-dimensional structure. As convolutional neural networks have a powerful ability to learn data features at multiple levels of abstraction, we deploy CNN to predict protein distance. To accelerate the training process, we apply a quantum convolutional neural network(QCNN) to improve the protein structure prediction efficiently. For QCNN, the conventional convolutional layer is transformed to a quantum convolution or quanvolutional layer. Since the protein data has large input, we explore the large dimension of these quantum transformations. And in experiments, we compare the different number of layers in QCNN during the training phase. We found the QCNN is similar to CNN that is the deeper layer can get better performance. The simulations show the proposed method can accelerate the convergence while maintaining the performance. Zhenhou Hong, Jianzong Wang, Xiaoyang Qu, Xinghua Zhu, Jing Xiao 0006 |
IJCNN | 3 |
| 2021 | When Hearing the Voice, Who Will Come to Your MindabstractSpeech is a carrier containing rich biological information, such as speaker identity information including age, gender, race. In this paper, we explore the use of a self-supervised method to obtain speaker identity information from high-dimensional speech representations to generate face image. At the same time, considering that the biological information contained in the same piece of speech has different expression forms (such as images), we designed a cross-modal knowledge distillation method to transform the feature information from the visual domain to the speech domain. The feature vectors obtained through self-supervised learning and knowledge distillation are fed into a GAN-based generative model to obtain facial images containing speaker information. Subjective experiments show that our model can reach a well performance in the task of speaker identification. Experiments show that our proposed method can effectively establish the connection between different modalities and generate a face with rich biological information. Zhenhou Hong, Jianzong Wang, Xiaoyang Qu, Zihang Wei, Jing Xiao 0006 |
IJCNN | 5 |
| 2021 | Communication-Memory-Efficient Decentralized Learning For Audio RepresentationabstractSmartphones and wearable devices produce a wealth of audio data, which cannot be accumulated in a centralized repository for learning supervised models due to privacy and bandwidth limitation. Federated learning provides a solution for learning model from decentralized data. But conventionally, it assumes the availability of labeled samples, whereas on-device data are generally unlabeled. For solving these issues, in this paper we propose the self-supervised learning approach in a federated manner without moving the unlabeled audio data. We try the audio albert as the self-supervised model, which achieves comparable performance to other pre-trained model but with smaller model size. The federated self-supervised framework has tremendous communication cost during training, and the transformer architecture utilized in audio albert has the problem of memory footprint, which are practical in loT devices. To address the first issue, we propose the Gradient Compression and CSR Encoding (GCE) to reduce communication requires each round. Furthermore, we apply the reversible idea to the transformer, which does not need to store the activation in each layer thus reduce the memory footprint. Moreover, we evaluate the quality of the self-supervised pre-training model under the federated setting, and the model achieves considerable performance in the downstream tasks by fine-tuning. Leilai Li, Jianzong Wang, Xiaoyang Qu, Jing Xiao 0006 |
IJCNN | 3 |
| 2021 | Automatic Joint Optimization of Algorithm-Level Compression and Compiler-Based Acceleration with Reinforcement Learning for DNN in Edge DevicesabstractMore accurate machine learning models often require more memory cost and more software-hardware co-adaption efforts for deployments on resource-constrained devices. Model compression techniques and deep learning compiler are developed to reduce the memory cost and latency. However, current methods require tremendous engineering efforts to optimize the model manually. This paper introduces a jointly learning based framework to perform the compression task and the acceleration task simultaneously. The joint optimization method auto-tunes the algorithm-level compression and compiler-based acceleration with reinforcement learning. The experiment results demonstrate that we compress the model by a factor of 2 or 8, and accelerate the optimization up to 30 times using our learning framework. Jianzong Wang, Xiaoyang Qu, Zihang Wei, Jing Xiao 0006 |
IJCNN | 3 |
| 2021 | Enhancing Neural Architecture Search by Upgrading Weak Components
Xiaoyang Qu, Jianzong Wang, Jing Xiao 0006 |
IJCNN | 1 |
| 2021 | Contrastive Learning for improving End-to-end Speaker VerificationabstractSpeaker verification involves examining the speech signal to authenticate the claim of a speaker as true or false. Deep neural networks are one of the successful implementations of complex non-linear models to learn unique and invariant features of data. They have been employed in speech recognition tasks and have shown their potential to be used for speaker recognition also. However, the overfitting problem is remained to prevent the model's performance. In this study, we apply contrastive learning on speaker verification tasks to solve the robustness problem. Besides, we introduce domain adaptive loss on the tasks. Experimental results and ablation study that indicate that our proposed model outperforms various baseline end-to-end methods significantly by at least relative 10%, including d-vector approaches, deep-speaker, and generalized end-to-end model, for text-dependent speaker verification on a company's internal text-dependent voice command DataSet. Yanxi Tang, Jianzong Wang, Xiaoyang Qu, Jing Xiao 0006 |
IJCNN | 3 |
| 2021 | CACnet: Cube Attentional CNN for Automatic Speech RecognitionabstractEnd-to-end models have been widely used in Automatic Speech Recognition (ASR). Convolutional Neural Networks (CNNs) can effectively use spectrum information to model acoustic models. However, the convolution layers have limitations on the receptive field leading to restrictions for long speech signals. Inspired by this, we propose a Cube Attention CNN network(CACnet) that uses two different attention blocks to integrate the feature information of different dimensions for extending context information. Thereinto, the Global Deep Attention Block utilizes non-local operations to compute interactions between any two positions on feature maps and enables the acquirement of global feature representations while the Cross-Channel Attention Block adaptively recalibrates channel-wise feature responses. Then, outputs of the above two attention modules will be added up to further improve the feature representation which contributes to enrich contextual information. Finally, the performance of our proposed architecture will be explored under ASR tasks in English circumstances. Experiments on LibriSpeech indicate that CACnet achieves a word error rate (WER) of 3.78%/9.56% without language model (LM), and 2.84%/6.97% with LM, which is near state-of-the-art accuracy. CACnet on WSJ with 4.4% WER obtains better performance, compared to CTC-based CNN models, such as QuartzNet and Jasper, with the same language model. The proposed network achieves competitive accuracy while having fewer parameters. Moreover, CACnet can be easily incorporated into any existed network since it has the same input and output dimensions. Jianzong Wang, Xiaoyang Qu, Ning Cheng 0001, Jing Xiao 0006 |
IJCNN | 4 |
| 2021 | Federated Learning with Dynamic Transformer for Text to SpeechabstractText to speech (TTS) is a crucial task for user interaction, but TTS model training relies on a sizable set of high-quality original datasets. Due to privacy and security issues, the original datasets are usually unavailable directly. Recently, federated learning proposes a popular distributed machine learning paradigm with an enhanced privacy protection mechanism. It offers a practical and secure framework for data owners to collaborate with others, thus obtaining a better global model trained on the larger dataset. However, due to the high complexity of transformer models, the convergence process becomes slow and unstable in the federated learning setting. Besides, the transformer model trained in federated learning is costly communication and limited computational speed on clients, impeding its popularity. To deal with these challenges, we propose the federated dynamic transformer. On the one hand, the performance is greatly improved comparing with the federated transformer, approaching centralize-trained Transformer-TTS when increasing clients number. On the other hand, it achieves faster and more stable convergence in the training phase and significantly reduces communication time. Experiments on the LJSpeech dataset also strongly prove our method's advantage. Zhenhou Hong, Jianzong Wang, Xiaoyang Qu, Chendong Zhao, Jing Xiao 0006 |
Interspeech | 3 |
| 2021 | Effective Phase Encoding for End-To-End Speaker Verification
Junyi Peng, Xiaoyang Qu, Rongzhi Gu, Jianzong Wang, Jing Xiao 0006, Lukás Burget, Jan Cernocký |
Interspeech | 2 |
| 2021 | ICSpk: Interpretable Complex Speaker Embedding Extractor from Raw Waveform
Junyi Peng, Xiaoyang Qu, Jianzong Wang, Rongzhi Gu, Jing Xiao 0006, Lukás Burget, Jan Cernocký |
Interspeech | 2 |
| 2021 | Speech2Video: Cross-Modal Distillation for Speech to Video GenerationabstractThis paper investigates a novel task of talking face video generation solely from speeches.The speech-to-video generation technique can spark interesting applications in entertainment, customer service, and human-computer-interaction industries.Indeed, the timbre, accent and speed in speeches could contain rich information relevant to speakers' appearance.The challenge mainly lies in disentangling the distinct visual attributes from audio signals.In this article, we propose a light-weight, cross-modal distillation method to extract disentangled emotional and identity information from unlabelled video inputs.The extracted features are then integrated by a generative adversarial network into talking face video clips.With carefully crafted discriminators, the proposed framework achieves realistic generation results.Experiments with observed individuals demonstrated that the proposed framework captures the emotional expressions solely from speeches, and produces spontaneous facial motion in the video output.Compared to the baseline method where speeches are combined with a static image of the speaker, the results of the proposed framework is almost indistinguishable.User studies also show that the proposed method outperforms the existing algorithms in terms of emotion expression in the generated videos. Shijing Si, Jianzong Wang, Xiaoyang Qu, Ning Cheng 0001, Xinghua Zhu, Jing Xiao 0006 |
Interspeech | 3 |
| 2021 | Variational Information Bottleneck for Effective Low-Resource Audio ClassificationabstractLarge-scale deep neural networks (DNNs) such as convolutional neural networks (CNNs) have achieved impressive performance in audio classification for their powerful capacity and strong generalization ability. However, when training a DNN model on low-resource tasks, it is usually prone to overfitting the small data and learning too much redundant information. To address this issue, we propose to use variational information bottleneck (VIB) to mitigate overfitting and suppress irrelevant information. In this work, we conduct experiments on a 4-layer CNN. However, the VIB framework is ready-to-use and could be easily utilized with many other state-of-the-art network architectures. Evaluation on a few audio datasets shows that our approach significantly outperforms baseline methods, yielding _ 5:0% improvement in terms of classification accuracy in some low-source settings. Copyright © 2021 ISCA. Shijing Si, Jianzong Wang, Huiming Sun, Jianhan Wu 0001, Chuanyao Zhang, Xiaoyang Qu, Ning Cheng 0001, Jing Xiao 0006 |
Interspeech | 6 |
| 2021 | Case Study of Few-Shot Learning in Text Recognition Models
Jianzong Wang, Shijing Si, Zhenhou Hong, Xiaoyang Qu, Xinghua Zhu, Jing Xiao 0006 |
WISE (2) | 4 |
| 2020 | Multi-objective Cuckoo Algorithm for Mobile Devices Network Architecture Search
Jianzong Wang, Xiaoyang Qu, Jing Xiao 0006 |
ICANN (1) | 4 |
| 2020 | Evolutionary Algorithm Enhanced Neural Architecture Search for Text-Independent Speaker VerificationabstractState-of-the-art speaker verification models are based on deep learning techniques, which heavily depend on the handdesigned neural architectures from experts or engineers. We borrow the idea of neural architecture search(NAS) for the textindependent speaker verification task. As NAS can learn deep network structures automatically, we introduce the NAS conception into the well-known x-vector network. Furthermore, this paper proposes an evolutionary algorithm enhanced neural architecture search method called Auto-Vector to automatically discover promising networks for the speaker verification task. The experimental results demonstrate our NAS-based model outperforms state-of-the-art speaker verification models. Xiaoyang Qu, Jianzong Wang, Jing Xiao 0006 |
INTERSPEECH | 1 |
| 2020 | 3D Point Cloud Segmentation for Complex Structure Based on PointSIFT
Jianzong Wang, Xiaoyang Qu, Jing Xiao 0006 |
PRCV (1) | 3 |
| 2018 | Workload Scheduling for Massive Storage Systems with Arbitrary Renewable SupplyabstractAs datacenters grow in scale, increasing energy costs and carbon emissions have led data centers to seek renewable energy, such as wind and solar energy. However, tackling the challenges associated with the intermittency and variability of renewable energy is difficult. This paper proposes a scheme called GreenMatch, which deploys an SSD cache to match green energy supplies with a time-shifting workload schedule while maintaining low latency for online data-intensive services. With the SSD cache, the process for a latency-sensitive request to access a disk is divided into two stages: a low-energy/low-latency online stage and a high-energy/high-latency off-line stage. As the process in the latter stage is off-line, it offers opportunities for time-shifting workload scheduling in response to variations of green energy supplies. We also allocate an HDD cache to guarantee data availability when renewable energy is inadequate. Furthermore, we design a novel replacement policy called Inactive P-disk First for the HDD cache to avoid inactive disk accesses. The experimental results show that GreenMatch can make full use of renewable energy while minimizing the negative impacts of intermittency and variability on performance and availability. Daping Li, Xiaoyang Qu, Jiguang Wan 0001, Jun Wang 0001, Xiaozhao Zhuang, Changsheng Xie 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2017 | OptiMatch: Enabling an Optimal Match between Green Power and Various Workloads for Renewable-Energy Powered Storage SystemsabstractTo reduce energy consumption and carbon emission, many data centers have deployed (or anticipate to build) their own renewable-energy power plants. However, the renewable energy (such as wind, tide, and solar energy) has the serious issues of intermittency and variability that prevent the green energy from being utilized effectively in practice. To cope with the issues, new power-supply management policies and workload scheduling algorithms have been designed. However, most existing work focuses on power optimization on computation only. In this paper, we introduce a novel scheme called OptiMatch to optimize the match between the power supply and the user-workload demand for massive storage systems that are mostly powered by renewable energy sources. OptiMatch has a hierarchical architecture, which consists of a number of heterogeneous storage devices. OptiMatch systematically utilizes the performance disparities between heterogeneous storage devices (i.e., performance per watt, IOPS/watt) to split the process for every write request into two stages: an on-line stage and a deferred off-line stage. The deferred off-line requests are used to match the green energy supplies. To maximize green energy utilization and minimize power budget without sacrificing quality of service, the fundamental methodology is to make the aggregate power supplies be proportional to the I/O workload demand at any time. To this end, our OptiMatch employs novel co-design optimizations. (1) We propose a dual-drive power control approach that makes the number of active nodes proportional to the workload demand when the green power supply is insufficient, meanwhile be proportional to the green power supply when green power is sufficient. (2) During periods of insufficient green supplies, we exploit virtualization consolidation schemes which enable a fine-grained power control to minimize the grid budgets. (3) During the periods of sufficient green supplies, we design an intelligent workload scheduling scheme which enables a near-optimal off-line requests assignment to maximize the green utilization. The experimental results demonstrate that the new OptiMatch framework can achieve high green utilization (up to 94.9%) with a minor performance degradation (less than 9.8%). Xiaoyang Qu, Jiguang Wan 0001, Fengguang Song, Xiaozhao Zhuang, Fei Wu 0005, Changsheng Xie 0001 |
ICPP | 1 |
| 2017 | DEFT-Cache: A Cost-Effective and Highly Reliable SSD Cache for RAID StorageabstractThis paper proposes a new SSD cache architecture, DEFT-cache, Delayed Erasing and Fast Taping, that maximizes I/O performance and reliability of RAID storage. First of all, DEFT-Cache exploits the inherent physical properties of flash memory SSD by making use of old data that have been overwritten but still in existence in SSD to minimize small write penalty of RAID5/6. As data pages being overwritten in SSD, old data pages are invalidated and become candidates for erasure and garbage collections. Our idea is to selectively delay the erasure of the pages and let these otherwise useless old data in SSD contribute to I/O performance for parity computations upon write I/Os. Secondly, DEFT-Cache provides inexpensive redundancy to the SSD cache by having one physical SSD and one virtual SSD as a mirror cache. The virtual SSD is implemented on HDD but using log-structured data layout, i.e. write data are quickly logged to HDD using sequential write. The dual and redundant caches provide a cost-effective and highly reliable write-back SSD cache. We have implemented DEFT-Cache on Linux system. Extensive experiments have been carried out to evaluate the potential benefits of our new techniques. Experimental results on SPC and Microsoft traces have shown that DEFT-Cache improves I/O performance by 26.81% to 56.26% in terms of average user response time. The virtual SSD mirror cache can absorb write I/Os as fast as physical SSD providing the same reliability as two physical SSD caches without noticeable performance loss. Jiguang Wan 0001, Qing Yang 0001, Xiaoyang Qu, Changsheng Xie 0001 |
IPDPS | 5 |
| 2017 | Exploiting Virtual Metadata Servers to Provide Multi-Level Consistency for Key-Value Object-Based Data StoreabstractDistributed data store is a fundamental building block for various Internet services. For large-scale distributed data store, the scalability and consistency of metadata services are prone to be the bottleneck. Various schemes are proposed to tackle the challenge of scalability and consistency within metadata services. While centralized single-node metadata services with low scalability provide low- overhead consistency maintenance, distributed metadata servers with high scalability often suffer complicated management and high-overhead consistency maintenance. As some key-value object-based storage systems locate and access an object by hashing function (e.g., consistent hashing table), there are no dedicated physical servers for metadata services. For key-value store without dedicated metadata servers, we exploited a scheme called virtual metadata servers (virtual MDS), which can create an opportunity to provide high performance and multi- level consistency. While conventional key-value data store distributes metadata across data nodes, our scheme uses proxy nodes, where virtual disks created, as virtual MDS to hold the metadata of virtual disks. Meanwhile, we also combine the characteristic of virtual disks and metadata services to implement a multi-level consistency strategy for the key-value object-based store without dedicated physical metadata servers. With virtual MDS, we use version information to update data asynchronously and check the version consistency periodically, then correct the stale entries properly. In this way, our virtual MDS can provide multi-level of consistency to cope with different read performance demand from users. The experiment results demonstrate that our scheme with relaxed consistency can enhance random write performance by 50% and improve random read performance by 16% compared with the standard storage system with strict consistency. Xiaozhao Zhuang, Xiaoyang Qu, Zhiyong Lu, Jiguang Wan 0001, Changsheng Xie 0001 |
NAS | 2 |
| 2017 | A reliable and energy-efficient storage system with erasure coding cacheabstractIn modern energy-saving replication storage systems, a primary group of disks is always powered up to serve incoming requests while other disks are often spun down to save energy during slack periods. However, since new writes cannot be immediately synchronized into all disks, system reliability is degraded. In this paper, we develop a high-reliability and energy-efficient replication storage system, named RERAID, based on RAID10. RERAID employs part of the free space in the primary disk group and uses erasure coding to construct a code cache at the front end to absorb new writes. Since code cache supports failure recovery of two or more disks by using erasure coding, RERAID guarantees a reliability comparable with that of the RAID10 storage system. In addition, we develop an algorithm, called erasure coding write (ECW), to buffer many small random writes into a few large writes, which are then written to the code cache in a parallel fashion sequentially to improve the write performance. Experimental results show that RERAID significantly improves write performance and saves more energy than existing solutions. Jiguang Wan 0001, Daping Li, Xiaoyang Qu, Jun Wang 0001, Changsheng Xie 0001 |
Frontiers Inf. Technol. Electron. Eng. | 3 |
| 2016 | GreenMatch: Renewable-Aware Workload Scheduling for Massive Storage SystemsabstractAs datacenters grow in scale, increasing energy costs and carbon emissions have led data centers to seek renewable energy, such as wind and solar energy. However, tackling the challenges associated with the intermittent nature and variability of renewable energy is substantial. This paper proposes a scheme called GreenMatch, which deploys an SSD-cache to match green energy supplies with a time-shifting workload schedule while maintaining low latency for online data-intensive services. With the SSD-cache, the process for a latency-sensitive request to access a disk is divided into two stages: a low-energy low-latency online stage and a high-energy high-latency off-line stage. As the process in the latter stage is off-line, it offers opportunities for time-shifting workload scheduling in response to variations of green energy supplies. We also allocate an HDD-cache to guarantee data availability when renewable energy is non-adequate. Furthermore, we design a novel replacement policy called Inactive Disk First for the HDD-cache to avoid inactive disk accesses. The experimental results show that GreenMatch can make full use of renewable energy while minimizing the negative impact of intermittency and variability on performance and availability. Xiaoyang Qu, Jiguang Wan 0001, Jun Wang 0001, Liqiong Liu, Changsheng Xie 0001 |
IPDPS | 1 |
| 2016 | CircularCache: Scalable and Adaptive Cache Management for Massive Storage SystemsabstractIn order to enhance the performance of HDD-based storage systems, low-latency and high-IOPS SSDs are usually deployed as a cache above HDDs. With explosive data growth, a large-scale SSD-based cache tend to adopt partition management for overall cached data distribution across multiple cache nodes. We proposed an adaptive and scalable SSD- based cache called CircularCache, which distributes hot data across multiple cache nodes. The hotter virtual disks deserve more allocated free space in the SSD-cache. This paper exploited a dynamic replacement algorithm called VBQ(VDI-Based Queues) to manage the SSD-cache. The VBQ scheme manages the SSD-cache by dynamically manipulating the upper- bounds and lower-bounds of multiple queues based on the total access number of virtual disks. To mitigate negative impacts of destaging on overall storage performance, the dirty data in the cache will be written back to data nodes during idle time. At the same time, we utilize the redundant storage space in the data nodes as logging area to retain reliability of the dirty data on the SSDcache. The prototype of CircularCache is implemented based on Sheepdog. Experimental results show that CircularCache offers a performance improvement by up to 270% compared with the standard distributed storage system without an SSD-based cache. Liqiong Liu, Xiaoyang Qu, Yubiao Zhang, Xiaodong Yi 0003, Siwang Zeng, Jiguang Wan 0001, Changsheng Xie 0001 |
NAS | 2 |
| 2016 | DVS: Dynamic Variable-Width Striping RAID for Shingled Write DisksabstractDisk data density improvement will eventually be limited by the super-paramagnetic effect for perpendicular magnetic recording. Of the various new technologies being explored, Shingled Magnetic Recording (SMR) exposes as the most promising one to achieve high areal density and only make little changes to the manufacturing process. At present, high-capacity SMR drives are available from Seagate and HGST. Since SMR is leading next generation disk technology and increasing SMR drives will be used in storage systems, there is a great need to look over the current RAID storage techniques based on HDDs again. In this paper, we proposed a dynamic variable-width striping RAID (DVS-RAID) for SMR drives to reduce the parity updating cost. DVS-RAID never overwrites the old data, but always constructs a new full or partial stripe (variable-width stripe), and writes to the SMR drives through appending. In addition, taking the access characteristics of SMR drives into consideration, we present a new write cache management that exploits both spatial and temporal localities. The experiment with six real-world traces demonstrates that DVS- RAID exhibits a slightly lower performance than HDD- based RAID on update intensive workloads. However the performance of DVS-RAID is better than HDD-based RAID with sequential access, read-dominated workloads or workloads with rarely update. Ting Yao 0001, Xiaoyang Qu, Jiguang Wan 0001, Changsheng Xie 0001 |
NAS | 3 |
| 2015 | ThinRAID: Thinning Down RAID Array for Energy ConservationabstractThe current power managements in RAID array are mostly designed to conserve energy by spinning down partial disks of standard RAID architecture. However, spinning down several disks not only decreases disk parallelism, but also creates new problems, for example, partial chunks of the stripe cannot be accessed directly or multiple chunks of the same stripe are stored on the same disk, which affect spatial locality. We refer these problems as stripe degradation, which results in further performance degradation. To avoid such problems, this paper proposes a new RAID storage architecture called ThinRAID, which uses a subset of disks to build a capacity-adaptive RAID array based on the volume of the data set. Also, the other non-essential disks are spun down to save energy. When the workload is projected to become heavier based on our forecast model, data are migrated to disks that have recently transitioned from standby to active. Furthermore, we also propose a novel data reorganization algorithm that can minimize data migration. We have implemented ThinRAID in the Linux kernel and evaluated its performance and energy efficiency by replaying seven representative traces. Experimental results show that ThinRAID can save 15-27 percent on energy on average over conventional RAID, with minimum performance degradation. In comparison to PARAID, ThinRAID achieves up to 62 percent performance improvement. Jiguang Wan 0001, Xiaoyang Qu, Jun Wang 0001, Changsheng Xie 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |