Xuan Wang 0002

dblp:34/4799-2 · DBLP profile ↗
← Back
166ranked-venue papers
5as first author
85since 2021 · last 2026
0000-0002-3512-0649ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 56 · 31 since 2021Applied, interdisciplinary, general and emerging computing · 35 · 5 first-author · 17 since 2021Systems, architecture and hardware · 23 · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 23 · 11 since 2021Databases, data management, data science and information retrieval · 14 · 12 since 2021Security and privacy · 12 · 4 since 2021Computer networks · 10 · 4 since 2021Human-computer interaction and ubiquitous computing · 10 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 4 · 4 since 2021
YearPublicationVenuePosition
2026 BulletTime4D: Towards High Spatio-Temporal Resolution Dynamic Scene Rendering via Spike-Guided Stereo Vision
abstract
High spatio‑temporal resolution novel‑view scene rendering is crucial for applications such as sports analysis and scientific experiments. However, existing Dynamic Scene Rendering (DSR) approaches typically rely on conventional RGB cameras with limited frame rates, making it difficult to achieve high spatio‑temporal resolution. In this paper, we present BulletTime4D, a high spatio‑temporal resolution DSR framework, which is the first trial to integrate a spike camera with binocular RGB cameras for dynamic scene reconstruction. Specifically, we first develop a hybrid camera prototype and build a real‑world dynamic scene reconstruction dataset. Then, BulletTime4D presents a multi‑timescale deformation representation by combining low‑frequency spatio‑temporal features with high‑frequency inter‑frame motion features. Finally, a rendering network is designed capable of projecting 4D Gaussians into the spike domain for spike rendering, and a cross‑domain supervision strategy is proposed to achieve high‑frame‑rate texture and color rendering. The results show that BulletTime4D outperforms state‑of‑the‑art methods on both simulated and real‑world datasets. In addition, BulletTime4D can synthesize 300 FPS novel‑view renderings using stereo RGB cameras at 30 FPS and a single spike camera.
Yiqian Chang, Haoran Xu 0004, Qinghong Ye, Jianing Li 0001, Xuan Wang 0002, Wei Zhang 0161, Peixi Peng
AAAI5
2026 BDI-based Opponent Modeling and Strategy Generation for Multi-Issue Negotiation (Student Abstract)
abstract
Accurately modeling opponent behaviors and integrating strategy are key challenges for multi-issue automated negotiation. Existing approaches often isolate preference learning or trend prediction and lack a unified cognitive structure with coordinated reasoning. This paper proposes a BDI (Belief-Desire-Intention)-based opponent modeling and strategy generation framework. The framework analyzes opponent responses (Belief), predicts preference weights and the utility function (Desire), and infers utilities of future offers (Intention). Building on these predictions, we design a responsive strategy, enabling gradual concessions and balanced outcomes. Our main contributions are: D-MBUE in the Desire module, I-DABI in the Intention module, and the BDI Negotiator on top of the modeling modules. Experiments on 45 standard negotiation domains and against 12 representative opponents demonstrate the effectiveness of our BDI framework.
Tianzi Ma, Yulin Wu 0001, Hang Ren 0002, Xiaozhen Sun, Shuhan Qi, Xuan Wang 0002
AAAI6
2026 NegLLM: Enhancing Strategic LLM Negotiation Agents with Case-Based Reasoning
Ruoke Wang, Tianzi Ma, Yulin Wu 0001, Jiajia Zhang 0001, Xuan Wang 0002
ICCBR5
2026 Policy Extraction-Based Adversarial Attack in Multi-Agent Reinforcement Learning
Bang Zhang, Wenjian Luo, Kesheng Chen, Yujiang Liu, Shuhan Qi, Xuan Wang 0002
ICIC (2)6
2026 HierMoE: Accelerating MoE Training with Hierarchical Token Deduplication and Expert Swap
Wenxiang Lin, Xinglin Pan, Lin Zhang 0059, Shaohuai Shi, Xuan Wang 0002, Xiaowen Chu 0001
INFOCOM5
2026 CryoDETR: a Deformable DETR-Based Method for Particle Picking in Cryo-EM Micrographs
Xuan Wang 0002, Fa Zhang 0001
ISBRA (2)1
2026 AWMA-MoE: Attention-Guided Watermark Adapter with MoE for Latent Diffusion Models
abstract
With the evolving generative models, generated images are closer to reality, raising concerns about information authenticity and malicious misuse. Invisible watermarks offer a practical approach to detecting and tracing them. However, while image watermarking inevitably introduces quality degradation, most existing methods primarily focus on improving watermark robustness. To address this limitation, we propose AWMA-MoE, a framework that enhances the quality of generated images while preserving strong watermark robustness. Specifically, we design an attention-based adapter that adaptively embeds watermarks with spatially varying strengths across image regions. Building upon this, we introduce an MoE architecture that leverages diverse experts to further improve image quality while retaining watermark robustness. Experiments demonstrate that AWMA-MoE can reduce the distortion of generated images and exhibit competitive watermark performance, thus striking an improved balance for watermarking generated image tasks and better linking post-hoc and in-generation methods.
Xinyu Xiao, Jian Zhang 0019, Shuhan Qi, Yulin Wu 0001, Xuan Wang 0002
WWW6
2026 SPVR: syntax-to-prompt vulnerability repair based on large language models
Ruoke Wang, Zongjie Li, Cuiyun Gao 0001, Chaozheng Wang, Yang Xiao 0011, Xuan Wang 0002
Autom. Softw. Eng.6
2026 MZSGO: multimodal zero-shot protein function annotation via evolutionary signals and textual semantics
abstract
MOTIVATION: Although deep learning has significantly advanced the field of protein function prediction, current approaches are limited by their reliance on a narrow set of modalities. Specifically, they primarily rely on sequence patterns and treat protein domain data and functional labels merely as categorical tags. Consequently, they fail to capitalize on the semantic richness embedded within their textual definitions. These constraints hinder their ability to generalize to novel labels. To tackle this issue, we present MZSGO, a multimodal zero-shot framework that fuses evolutionary signals from protein language models with semantic features derived from large language models (LLMs). By employing an adaptive gated fusion mechanism, MZSGO effectively aligns sequence-based and text-based modalities to enable robust predictions for unseen labels. RESULTS: By unifying protein representations and functional annotations, we bridge the semantic gap that limits current approaches. Results indicate that while our model remains competitive on supervised benchmarks, it demonstrates a marked advantage over existing methods in zero-shot tasks. It specifically excels at recognizing previously unseen long-tail and novel Gene Ontology (GO) terms. AVAILABILITY AND IMPLEMENTATION: The source code and datasets are available at https://github.com/toxic-byte/MZSGO.
Boyue Cui, Yujuan Li, Shiqu Chen, Jiaming Wei, Xuan Wang 0002, Yadong Wang 0001, Junyi Li 0004
Bioinform.5
2026 GR2ST: spatial transcriptomics prediction based on graph-enhanced multimodal contrastive learning
abstract
MOTIVATION: Spatial transcriptomics techniques capture gene expression data and spatial coordinates, while simultaneously correlating them with tissue section images. This advantage makes Spatial transcriptomics data highly valuable for research, such as investigating disease mechanisms and cancer prognosis. However, the extended time and high cost of spatial transcriptomic sequencing currently limit further advancements in this field. The development of numerous deep learning methods aimed at predicting spatial transcriptomics from histology images has advanced significantly. However, these approaches often lack the ability to effectively integrate histology images with spatial transcriptomic data. Here, we propose GR2ST, a deep learning model that learns the underlying connections between image features and gene expression to predict spatial transcriptomics. RESULTS: GR2ST leverages a large pre-trained pathology model to extract high-level histological features. We designed a dual-branch graph architecture, consisting of a dynamic threshold-based functional graph and a radius-constrained spatial graph, to capture complex spot interactions within heterogeneous tissues. The model aligns histology images with gene expression representations through a multimodal contrastive learning framework. It achieves adaptive gene expression generation via a Cell-Type Guided Multi-Branch Regression Head supervised by a context-aware weighting network, which is further integrated with cross-sample retrieval to construct an ensemble prediction. The performance of the model is evaluated on three cancer-related spatial transcriptomics datasets, including cutaneous squamous cell carcinoma and two human breast cancer cohorts, to demonstrate its effectiveness and robustness. AVAILABILITY: https://github.com/zjl1109294570/GR2ST.
Jingli Zhou, Xuan Wang 0002, Yadong Wang 0001, Junyi Li 0004
Bioinform.4
2026 Solving equilibrium for adversarial team games utilizing fictitious team play with refined team plans
Jinheng Xiao, Chen Qiu 0003, Jiajia Zhang 0001, Shuhan Qi, Xuan Wang 0002
Expert Syst. Appl.6
2026 Distributed scalable multi-agent reinforcement learning with intrinsic-episodic dual exploration
Shuhan Qi, Shuhao Zhang 0009, Qiang Wang 0022, Jiajia Zhang 0001, Xuan Wang 0002
Future Gener. Comput. Syst.5
2026 Uncertainty-aware mixture of experts for robust multimodal sentiment analysis
Xinyu Xiao, Xuan Wang 0002, Shuhan Qi, Huale Li
Pattern Recognit.2
2026 Leveraging Neural Architecture Search for improved downstream-agnostic adversarial attack
Haodong Xiao, Bin Chen 0011, Hao Fang 0011, Yulin Wu 0001, Xuan Wang 0002, Zhi Wang 0001, Shutao Xia
Pattern Recognit.7
2026 Orchestrating Prompt Expertise: Enhancing Knowledge Distillation via Expert-Guided Tuning
abstract
Multi-teacher knowledge distillation transfers knowledge from multiple large teacher models to a small student model and has performed well on many downstream tasks. However, when distilling knowledge from multiple teachers, it always suffers from the severe problems of being time-consuming and storage-extensive for multiple teacher models training and inference. We present MoE-KD, a simple but effective framework that produces supervision for training the student model from one single teacher model, which fixes the above problems and improves effectiveness. In the proposed MoE-KD, multiple trainable prompts are used to extract different views of samples from a single pre-trained language model and only a few parameters (prompts) need to be trained and stored. To guarantee the generated supervision signals with increased robustness and correctness, we introduce an uncertainty-based mechanism and a selector module, which routes the input instance to its corresponding teacher. We have also extended MoE KD to lifelong learning scenarios, proposing a lightweight solution for catastrophic forgetting. We conduct experiments on traditional KD scenarios and lifelong learning scenarios. MoE-KD yields improvements up to 1.1% and 140% in accuracy and efficiency in knowledge distillation and 2.8% improvements on average in lifelong learning, compared with the strong baseline methods.
Jun Rao, Shuhan Qi, Xuebo Liu 0002, Lei Wang 0203, Min Zhang 0005, Xuan Wang 0002
ACM Trans. Asian Low Resour. Lang. Inf. Process.7
2026 SAH-NeRF: Enhancing NeRF on Novel View Synthesis With an SNN-ANN Hybrid Framework
abstract
Neural Radiance Field (NeRF) utilizes Artificial Neural Networks (ANNs) to map 3D points and their corresponding 2D viewing directions to colors and densities. However, this approach encounters several challenges due to the inherent characteristics of ANNs. Firstly, ANNs tend to extract smooth features across sampled points, which makes it difficult to accurately represent the non-smooth variations between object surfaces and the surrounding air. Secondly, ANNs compute densities and colors independently for each point, failing to account for the sequential dependencies among points along the same ray. To tackle these issues, Spiking Neural Networks (SNNs) are introduced, which are better at processing sequential information and non-smooth representations. In this work, we propose a novel hybrid NeRF framework called SAH-NeRF that combines ANNs with SNNs. By harnessing ANNs’ robust representation capabilities alongside SNNs’ strengths in handling non-smooth distributions, our method could significantly improve the performance of three ANN-based NeRFs, surpassing state-of-the-art methods including 3D Gaussian Splatting. Notably, our SAH-NeRF could meanwhile enhance novel view synthesis and reduce energy consumption.
Yiqian Chang, Peixi Peng, Zhaokun Zhou, Xuan Wang 0002, Yonghong Tian 0001
IEEE Trans. Circuits Syst. Video Technol.4
2026 Dual Feature Fusion for Incomplete Multi-View Multi-Label Learning
abstract
Multi-view Multi-label Learning (MVML) aims to leverage multi-view information from input samples to achieve accurate predictions of multiple labels. Unfortunately, most existing MVML methods operate under the assumption of data completeness, which makes them ineffective in practical scenarios involving missing views or uncertain labels. Recent methods address incomplete data, but few approaches handle scenarios where both views and labels are missing. To address this challenge, we propose a Dual-view Feature-guided Fusion Learning (DFFL) framework. DFFL considers both view-specific unique features and inter-view consistent features. Specifically, DFFL constructs view-uniqueness contrastive learning to ensure that features within the same view maintain high semantic relevance under the condition of view missing, while the semantics between different views are different. Unlike previous methods, DFFL assumes that label relevance can be reversely mapped to high-dimensional features. By establishing View-consistency learning, the mutual information in the shared embedding space is maximized to achieve consistent feature alignment. In particular, DFFL minimizes the conditional entropy of the marginal distribution of multi-view features through dual prediction, thereby deriving the maximum joint distribution of feature fusion and combining the missing view index matrix to achieve feature fusion. This process can effectively alleviate the fusion feature suppression existing in previous methods. Finally, the missing label index matrix is combined with the fusion feature to complete the classification task. We validate the framework on five widely used datasets, and experimental results demonstrate that our approach achieves superior performance compared to state-of-the-art methods. Ablation studies further validated the effectiveness of each component in DFFL.
Xinyu Xiao, Shuhan Qi, Yulin Wu 0001, Bin Chen 0011, Xuan Wang 0002
IEEE Trans. Multim.6
2026 KPGS: Toward Real-World Complex Dynamic Scene Rendering With Keyframe-Driven Predictable Gaussian Splatting
abstract
Rendering complex dynamic scenes offers the advantage of observing and understanding the real world. However, existing Dynamic Scene Rendering (DSR) methods remain challenged by suboptimal reconstruction fidelity. These limitations stem from relying on a single, unified deformation model, which struggles to capture complex motions involving multiple sub-motions and abrupt geometric transitions. While temporal decomposition methods could alleviate such shortcomings, they introduce the additional challenge of ignoring motion correlations and increasing storage requirements. To address these issues, we introduce Keyframe-driven Predictable Gaussian Splatting (KPGS)-an efficient framework for high-fidelity complex dynamic scene rendering. First, we present a patch-wise HSV clustering for extracting keyframes. Second, a prediction network based on the Transformer is utilized to calculate the deformable Gaussians at discrete keyframe times via voxelization. Third, we propose an inter-frame deformation network and a mutual supervision between adjacent segments to maintain the temporal continuity. Extensive experiments on our newly built dataset (MotionGS), as well as public benchmarks HyperNeRF and Neu3D, demonstrate that KPGS could achieve a higher average view synthesis performance than SOTA approaches, while maintaining a balance between storage cost and performance. More details of the demo and dataset are available at KPGS Supplementary.
Yiqian Chang, Haoran Xu 0004, Jianing Li 0001, Xuan Wang 0002, Yonghong Tian 0001, Peixi Peng
IEEE Trans. Vis. Comput. Graph.4
2025 Towards Building Human-like Smart Agents in Modern 3D Video Games (Student Abstract)
abstract
In recent years, reinforcement learning has been widely applied in the field of games. However, most studies focus on assisting agents to achieve victory, with less attention paid to whether the agents exhibit human-like characteristics. In order to build human-like agents with high performance, we propose a method for learning the strategies of human players in modern three-dimensional video games. Our method utilizes a hierarchical framework, learning basic behaviors and intentions of human players at the lower level through imitation learning, and generalized policies at the high level through reinforcement learning. Compared with other existing methods, our method demonstrates significant advantages in learning human-like strategies in complex environments.
Zhihang Sun, Shuhan Qi, Xinhao Huang, Xinyu Xiao, Jiajia Zhang 0001, Xuan Wang 0002, Peixi Peng
AAAI6
2025 EMST: An Interpretable Multi-Modal Model for Spatial Transcriptomic Data Analysis
abstract
Spatial transcriptomic data provides gene expression, spatial positions, and histopathological images for each spot. However, due to technical difficulties in obtaining high-quality spatial transcriptomic data, it is crucial to enhance the quality of raw gene expression data through computational methods. Existing methods can not effectively fuse gene features and morphological features, and deep learning-based methods lack interpretability. Based on this, we propose EMST, an interpretable multi-modal model for spatial transcriptomic data analysis. We first extract morphological features of spots using a pre-trained large biological model, and then integrate morphological and gene features through a co-attention-based multi-modal fusion module. The fused features undergo an interpretable contrastive learning model to complete feature learning and enhance original gene expression data. Comparative experimental results on multiple datasets show that EMST can effectively enhance gene expression. We also conducted an interpretability analysis to make the training process of the model clearer.
Weiwei Yuan, Xuan Wang 0002, Junyi Li 0004
BIBM3
2025 Efficient Constant-Size Linkable Ring Signatures for Ad-Hoc Rings via Pairing-Based Set Membership Arguments
abstract
Linkable Ring Signatures (LRS) allow users to anonymously sign messages on behalf of ad-hoc rings, while ensuring that multiple signatures from the same user can be linked. This feature makes LRS widely used in privacy-preserving applications like e-voting and e-cash. To scale to systems with large user groups, efficient schemes with short signatures and fast verification are essential. Recent works, such as DualDory (ESORICS'22) and LLRing (ESORICS'24), improve verification efficiency through offline precomputations but rely on static rings, limiting their applicability in ad-hoc ring scenarios. Similarly, constant-size ring signature schemes based on accumulators face the same limitation.
Zhengzhou Tu, Man Ho Au, Xuan Wang 0002, Zoe Lin Jiang
CCS5
2025 Last-Iterate Convergence in Adaptive Regret Minimization for Approximate Extensive-Form Perfect Equilibrium
abstract
The Nash Equilibrium (NE) assumes rational play in imperfect-information Extensive-Form Games (EFGs) but fails to ensure optimal strategies for off-equilibrium branches of the game tree, potentially leading to suboptimal outcomes in practical settings. To address this, the Extensive-Form Perfect Equilibrium (EFPE), a refinement of NE, introduces controlled perturbations to model potential player errors. However, existing EFPE-finding algorithms, which typically rely on average strategy convergence and fixed perturbations, face significant limitations: computing average strategies incurs high computational costs and approximation errors, while fixed perturbations create a trade-off between NE approximation accuracy and the convergence rate of NE refinements. To tackle these challenges, we propose an efficient adaptive regret minimization algorithm for computing approximate EFPE, achieving last-iterate convergence in two-player zero-sum EFGs. Our approach introduces Reward Transformation Counterfactual Regret Minimization (RTCFR) to solve perturbed games and defines a novel metric, the Information Set Nash Equilibrium (ISNE), to dynamically adjust perturbations. Theoretical analysis confirms convergence to EFPE, and experimental results demonstrate that our method significantly outperforms state-of-the-art algorithms in both NE and EFPE-finding tasks.
Hang Ren 0002, Xiaozhen Sun, Tianzi Ma, Jiajia Zhang 0001, Xuan Wang 0002
ECAI5
2025 Merge then Realign: Simple and Effective Modality-Incremental Continual Learning for Multimodal LLMs
abstract
Recent advances in Multimodal Large Language Models (MLLMs) have enhanced their versatility as they integrate a growing number of modalities.Considering the heavy cost of training MLLMs, it is efficient to reuse the existing ones and extend them to more modalities through Modality-incremental Continual Learning (MCL).The exploration of MCL is in its early stages.In this work, we dive into the causes of performance degradation in MCL.We uncover that it suffers not only from forgetting as in traditional continual learning, but also from misalignment between the modality-agnostic and modality-specific components.To this end, we propose an elegantly simple MCL paradigm called "MErge then ReAlign" (MERA) to address both forgetting and misalignment.MERA avoids introducing heavy model budgets or modifying model architectures, hence is easy to deploy and highly reusable in the MLLM community.Extensive experiments demonstrate the impressive performance of MERA, holding an average of 99.84% Backward Relative Gain when extending to four modalities, achieving nearly lossless MCL performance.Our findings underscore the misalignment issue in MCL.More broadly, our work showcases how to adjust different components of MLLMs during continual learning.
Dingkun Zhang, Shuhan Qi, Xinyu Xiao, Kehai Chen, Xuan Wang 0002
EMNLP5
2025 ScheInfer: Efficient Inference of Large Language Models with Task Scheduling on Moderate GPUs
Wenxiang Lin, Xinglin Pan, Shaohuai Shi, Xuan Wang 0002, Xiaowen Chu 0001
Euro-Par (3)4
2025 Mast: Efficient Training of Mixture-of-Experts Transformers with Task Pipelining and Ordering
abstract
The utilization of the sparsely activated mixture-of-experts (MoE) technique has enabled the expansion of modern large language models (LLMs) to trillion-level sizes while maintaining a sub-linear increase in computations. This involves equipping an MoE layer with multiple experts, where only one or two experts are activated for each input data. However, the dynamic activation of MoE experts introduces extensive communications, limiting the scaling efficiency of distributed systems. In this work, we propose Mast to efficiently train MoE models by pipelining and re-ordering communication and computation tasks to effectively hide communication costs. Specifically, we first propose to overlap tasks in both attention layers and MoE layers. Then we theoretically analyze the task overlaps between communications and computations, identifying the inefficiencies of existing schedules. We then develop an optimization formulation to determine a near-optimal order for task pipelining with the objective of minimizing iteration time. We conduct extensive experiments on two 32-GPU clusters employing 432 configured MoE layers and three real-world MoE models based on BERT, GPT-2 and Mistral. The experimental results demonstrate that Mast outperforms state-of-the-art MoE training systems (DeepSpeed-MoE, Tutel, PipeMoE and CoCoNet) with an average speedup 1.13 ×-1.43 × on the MoE models.
Wenxiang Lin, Xinglin Pan, Shaohuai Shi, Xuan Wang 0002, Bo Li 0001, Xiaowen Chu 0001
ICDCS4
2025 MKDTI: Predicting Drug-Target Interactions via Multiple Kernel Fusion on Graph Attention Network
Yuhuan Zhou, Yaqiu Wang, Yulin Wu 0001, Qian Chen 0028, Weiwei Yuan, Xuan Wang 0002, Junyi Li 0004
ICIC (25)6
2025 Black-Box Adversarial Robustness Testing with Partial Observation for Multi-Agent Reinforcement Learning
abstract
Multi-Agent Reinforcement Learning (MARL) has shown great success in many aspects. However, the cooperative policy trained by MARL is vulnerable to adversarial attacks towards agents' observations, which could cause immeasurable damage to the agent team. A few techniques have been developed to test the robustness of MARL, but the feasibility of real implementation is not carefully considered. In this work, we propose a two-step framework to conduct destructive and sparse attacks under realistic conditions. The first step contains two attack scenes: Ally Observation Attack (AOA) and Enemy Observation Attack (EOA), which have limitations on the use of agents' observations. The first step selects victim agent from the team and screens the attack time, while the second step generates adversarial perturbation and adds it to the observation of the chosen victim in a black-box environment. To the best of our knowledge, this is the first work to test the robustness with partial observation and conduct the test in a black-box environment. Experiments on SMAC environments demonstrate that our methods show great performance in reducing the win rate and the team reward of the agent team trained by QMIX algorithm with lower perturbation steps.
Bang Zhang, Wenjian Luo, Kesheng Chen, Yujiang Liu, Shuhan Qi, Xuan Wang 0002
ICPADS6
2025 Natural Language to Overpass Query: A Multi-Step Approach Using Task Decomposition and Key-Value Correction
abstract
We investigate the challenge of generating OverpassQL from natural language in the Text-to-OverpassQL task and explore the data in the existing OverpassNL dataset. To address the structural mismatch between natural language and OverpassQL, we propose a task decomposition-based multi-step prompting approach that generates auxiliary information to help align natural language with OverpassQL structures, thereby enhancing model performance. Furthermore, we introduce a Key-Value Correction Module specifically targeting key-value pair matching difficulties in Text-to-OverpassQL tasks, designed to rectify potential syntactic errors and key-value mismatches in generated queries. Our experiments on GPT-3.5 Turbo and GPT-4 demonstrate absolute performance gains of$\mathbf{1. 4 \%}$and 0.6 % respectively. Under retrieval-augmented setting ablation, we achieve a more significant 3.5 % improvement with GPT-3.5 Turbo. Experimental results confirm that our method consistently improves performance across various models and configurations, particularly showing enhanced effectiveness in medium and small-scale models.
Xinrui Zhu, Xuan Wang 0002, Yuanfeng Song, Hanlin Gu
MDM3
2025 Spike4DGS: Towards High-Speed Dynamic Scene Rendering with 4D Gaussian Splatting via a Spike Camera Array
abstract
Spike camera with high temporal resolution offers a new perspective on high-speed dynamic scene rendering. Most existing rendering methods rely on Neural Radiance Fields (NeRF) or 3D Gaussian Splatting (3DGS) for static scenes using a monocular spike camera. However, these methods struggle with dynamic motion, while a single camera suffers from limited spatial coverage, making it challenging to reconstruct fine details in high-speed scenes. To address these problems, we propose Spike4DGS, the first high-speed dynamic scene rendering framework with 4D Gaussian Splatting using spike camera arrays. Technically, we first build a multi-view spike camera array to validate our solution, then establish both synthetic and real-world multi-view spike-based reconstruction datasets. Then, we design a multi-view spike-based dense initialization module that obtains dense point clouds and camera poses from continuous spike streams. Finally, we propose a spike-pixel synergy constraint supervision to optimize Spike4DGS, incorporating both rendered image quality loss and dynamic spatiotemporal spike loss. The results show that our Spike4DGS outperforms state-of-the-art methods in terms of novel view rendering quality on both synthetic and real-world datasets. More details are available at https://github.com/Qinghongye/Spike4DGS.
Qinghong Ye, Yiqian Chang, Jianing Li 0001, Haoran Xu 0004, Xuan Wang 0002, Wei Zhang 0161, Yonghong Tian 0001, Peixi Peng
NeurIPS5
2025 Information-Theoretic Point Cloud Defense: Harnessing Conditional Mutual Information Against Adversarial Attacks
Xinhao Zhong, Shuoyang Sun, Jiaxin Hong, Bin Chen 0011, Teko Ranoka, Xuan Wang 0002, Shutao Xia
PRCV (10)7
2025 Privacy-Preserving Social Recommendation: Privacy Leakage and Countermeasure
Yuyue Chen, Peng Yang 0016, Zoe Lin Jiang, Xuan Wang 0002, Chuanyi Liu
RecSys6
2025 MSNGO: multi-species protein function annotation based on 3D protein structure and network propagation
abstract
MOTIVATION: In recent years, protein function prediction has broken through the bottleneck of sequence features, significantly improving prediction accuracy using high-precision protein structures predicted by AlphaFold2. While single-species protein function prediction methods have achieved remarkable success, multi-species approaches still face challenges such as difficulties in multi-source data integration and insufficient knowledge transfer between distantly-related species. How to integrate large-scale data and provide effective cross-species label propagation for species with sparse protein annotations remains a critical and unresolved challenge. To address this problem, we propose the MSNGO (Multi-species protein Structures and Network to predict GO terms) model, which integrates structural features and network propagation methods. Our validation shows that using structural features can significantly improve the accuracy of multi-species protein function prediction. RESULTS: We employ graph representation learning techniques to extract amino acid representations from protein structure contact maps and train a structural model using a graph convolution pooling module to derive protein-level structural features. After incorporating the sequence features from ESM-2, we apply a network propagation algorithm to aggregate information and update node representations within a heterogeneous network. The results demonstrate that MSNGO outperforms previous multi-species protein function prediction methods that rely on sequence features and protein-protein networks. AVAILABILITY AND IMPLEMENTATION: https://github.com/blingbell/MSNGO.
Boyue Cui, Shiqu Chen, Xuan Wang 0002, Yadong Wang 0001, Junyi Li 0004
Bioinform.4
2025 Optimizing strategy selection in hidden role games
Chen Qiu 0003, Jinheng Xiao, Jiajia Zhang 0001, Shuhan Qi, Xuan Wang 0002
Eng. Appl. Artif. Intell.6
2025 FDAAC-CR: Practical Delegatable Attribute-Based Anonymous Credentials With Fine-Grained Delegation Management and Chainable Revocation
Peichen Ju, Yanqi Zhao, Zoe Lin Jiang, Man Ho Au, Yong Yu 0002, Xuan Wang 0002
IEEE Trans. Inf. Forensics Secur.8
2025 The Design of a High-Performance Fine-Grained Deduplication Framework for Backup Storage
abstract
Fine-grained deduplication (also known as delta compression) can achieve a better deduplication ratio compared to chunk-level deduplication. This technique removes not only identical chunks but also reduces redundancies between similar but non-identical chunks. Nevertheless, it introduces considerable I/O overhead in deduplication and restore processes, hindering the performance of these two processes and rendering fine-grained deduplication less popular than chunk-level deduplication to date. In this paper, we explore various issues that lead to additional I/O overhead and tackle them using several techniques. Moreover, we introduce MeGA, which attains fine-grained deduplication/restore speed nearly equivalent to chunk-level deduplication while maintaining the significant deduplication ratio benefit of fine-grained deduplication. Specifically, MeGA employs (1) a backup-workflow-oriented delta selector and cache-centric resemblance detection to mitigate poor spatial/temporal locality in the deduplication process, and (2) a delta-friendly data layout and “Always-Forward-Reference” traversal to address poor spatial/temporal locality in the restore workflow. Evaluations on four datasets show that MeGA achieves a better performance than other fine-grained deduplication approaches. Specifically, MeGA significantly outperforms the traditional greedy approach, providing 10–46 times better backup speed and 30–105 times more efficient restore speed, all while preserving a high deduplication ratio.
Xiangyu Zou, Wen Xia, Philip Shilane, Haijun Zhang 0002, Xuan Wang 0002
IEEE Trans. Parallel Distributed Syst.5
2025 Image Compression for Resource-Constrained AIoT System With Compressed Sensing
abstract
In today’s big data era, a key requirement is to implement intelligent semantic analysis (such as image recognition) on data gathered from an extensive array of smart devices in Artificial Intelligence IoT (AIoT) scenarios, all of which is processed at central cloud service providers. Recent advancements in deep-learning-based image compression have fostered semantic compression between machines. However, the deployment of an overparameterized encoder on Internet of Things (IoT) devices remains a challenge due to their restricted computing and storage capabilities. To tackle this issue, we propose a novel approach named compressed sensing (CS)-based asymmetric semantic image compression (CS-ASIC), explicitly designed for resource-constrained AIoT systems. This asymmetric semantic compression scheme intends to surpass the limitations of IoT devices, thereby facilitating efficient semantic compression for machine vision tasks. CS-ASIC notably includes a lightweight front encoder founded on deep image CS techniques, which utilizes rich image priors to learn measurement matrices for sampling. In tandem, a deep iterative decoder is designed cooperatively with the linear encoder offloaded at the server to enhance image reconstruction and semantic analysis across various semantic analysis tasks. Furthermore, we introduce a groundbreaking lossy CS semantic rate-distortion theoretical framework that justifies a compromise in rate for extended semantic distortion. Extensive experimental results underscore the superiority of the proposed CS-ASIC concerning the signal-semantic rate-distortion tradeoff, and its lower encoding complexity over existing codecs in an AIoT simulation environment.
Bin Chen 0011, Yujun Huang, Han Qiu 0001, Shutao Xia, Wei Fei, Xuan Wang 0002, Meikang Qiu
IEEE Trans. Syst. Man Cybern. Syst.6
2024 SelfAlign: Achieving Subtomogram Alignment with Self-Supervised Deep Learning
abstract
Cryo-Electron Tomography (Cryo-ET) and subtomogram averaging (STA) have been instrumental in advancing the analysis of high-resolution structural biology, enabling detailed insights into macromolecular complexes. However, due to limitations in sample thickness and electronic metrology, there are inherent issues with missing wedge artifacts and low signal-to-noise ratio in Cryo-ET. Researchers use STA to align and average subtomograms to address these two issues. Traditional STA methods, reliant on cross-correlation, are computationally expensive and not scalable for large datasets. The emerging method of using deep learning for STA has low accuracy and unstable performance at low signal-to-noise ratios. To address these issues, we proposed SelfAlign, a self-supervised deep learning approach for subtomogram alignment. To improve alignment accuracy, we introduce a rotation and translation method effectively reducing translation errors. Further, we present a self-labeling mechanism optimized for end-to-end processes,thereby abolishing the need for manual labeling. Additionally, we design a concise and efficient loss function to uphold stable training in scenarios with low signal-to-noise ratios. We demonstrate the efficacy of SelfAlign using four datasets, showcasing its superior performance in terms of alignment accuracy compared to existing methods. SelfAlign offers a robust and scalable solution for subtomogram analysis.
Xuan Wang 0002, Haofan Cao, Fa Zhang 0001
BIBM1
2024 Innovative 3D CTF Correction Techniques for Cryo-ET of Individual Particles
abstract
Cryo-electron tomography (Cryo-ET) and subtomogram averaging techniques are highly effective in revealing high-resolution molecular structures. In this technique, accurate Contrast Transfer Function (CTF) correction is essential, as it profoundly affects image quality. This impact arises from various factors, including electron beam characteristics, sample thickness, diffraction phenomena, and lens aberrations. Traditional two-dimensional CTF(2D-CTF) correction methods become inadequate when Cryo-ET is increasingly applied to thicker samples and multi-angle imaging scenarios, since variations in defocus along the tilt axis lead to resolution deterioration. To address this challenge, a novel methodology, termed the Focus Defocus Changes Contrast Transfer Function (FDC-CTF), is put forward as a new type of three-dimensional Contrast Transfer Function (3D-CTF) method. This method geometrically determines the degree of particle defocus based on their coordinates within tomographic slices, effectively tackling the defocus gradient issues induced by sample thickness and tilting. Furthermore, a new workflow is proposed where, after obtaining the particle coordinates, each particle undergoes individual 3D-CTF correction, thereby enhancing the overall accuracy and efficiency of the process. We use two datasets to demonstrate the effectiveness of our method, showcasing its superior performance in alignment accuracy compared to existing methods. It provides a powerful solution for tomographic reconstruction.
Xuan Wang 0002, Fa Zhang 0001
BIBM1
2024 Performance Analysis and Optimizations of Matrix Multiplications on ARMv8 Processors
abstract
General matrix multiplication (GEMM) as a fundamental subroutine has been widely used in many applications like scientific computing, machine learning, etc. Although many studies are dedicated to optimizing its performance, they mainly focus on matrices with regular shapes or x86 platforms. The irregularly shaped matrices on GEMM running on modern ARMv8 processors are under-explored. In this paper, we provide a thorough performance analysis of the general block-panel multiplication (GEBP) kernel of GEMM that has irregular shapes. Based on our analysis, we propose a new GEMM algorithm named EPPA with three novel schemes to improve GEMM performance on ARMv8 processors: i) eliminating packing to reduce Ll cache contention, ii) avoiding data eviction and pre-fetching data to reduce the Ll cache miss penalty, and iii) an adaptive selection strategy of the above two and original schemes. We conduct extensive experiments with a large range of irregular matrices on three popular ARMv8 processors compared to seven state-of-the-art GEMM libraries. The experimental results show that our EPPA algorithm outperforms existing ones across workloads and processors and accelerates real-world applications.
Hucheng Liu, Shaohuai Shi, Xuan Wang 0002, Zoe Lin Jiang, Qian Chen 0028
DATE3
2024 FSSiBNN: FSS-Based Secure Binarized Neural Network Inference with Free Bitwidth Conversion
Peng Yang 0016, Zoe Lin Jiang, Jiehang Zhuang, Siu-Ming Yiu, Xuan Wang 0002
ESORICS (1)6
2024 MUSE-Net: Disentangling Multi-Periodicity for Traffic Flow Forecasting
abstract
Accurate forecasting of traffic flow plays a crucial role in building smart cities in the new era. Previous work has achieved success in learning inherent spatial and temporal patterns of traffic flow. However, existing works investigated the multiple periodicities (e.g., hourly, daily, and weekly) of traffic via entanglement learning, which has not yet dealt with distribution shift and interaction shift problems in traffic flow. In this paper, we propose a novel disentanglement learning network, called MUSE-Net, to tackle the limitations of entanglement learning by simultaneously factorizing the exclusiveness and interaction of multi-periodic patterns in traffic flow. Grounded in the theory of mutual information, we first learn and dis-entangle exclusive and interactive representations of traffics from multi-periodic patterns. Then, we utilize semantic-pushing and semantic-pulling regularizations to encourage the learned representations to be independent and informative. Moreover, we derive a lower bound estimator to tractably optimize the disentanglement problem with multiple variables and propose a joint training model for traffic forecasting. Extensive experimental results on several real-world traffic datasets demonstrate the effectiveness of the proposed framework. The code is available at: https://github.com/JianyangQin/MUSE-Net.
Jianyang Qin, Yan Jia 0001, Yongxin Tong, Heyan Chai 0001, Ye Ding 0002, Xuan Wang 0002, Binxing Fang, Qing Liao 0001
ICDE6
2024 Scheduling Deep Learning Jobs in Multi-Tenant GPU Clusters via Wise Resource Sharing
abstract
Deep learning (DL) has demonstrated significant success across diverse fields, leading to the construction of dedicated GPU accelerators within GPU clusters for high-quality training services. Efficient scheduler designs for such clusters are vital to reduce operational costs and enhance resource utilization. While recent schedulers have shown impressive performance in optimizing DL job performance and cluster utilization through periodic reallocation or selection of GPU resources, they also encounter challenges such as preemption and migration overhead, along with potential DL accuracy degradation. Nonetheless, few explore the potential benefits of GPU sharing to improve resource utilization and reduce job queuing times.Motivated by these insights, we present a job scheduling model allowing multiple jobs to share the same set of GPUs without altering job training settings. We introduce SJF-BSBF (shortest job first with best sharing benefit first), a straightforward yet effective heuristic scheduling algorithm. SJF-BSBF intelligently selects job pairs for GPU resource sharing and runtime settings (sub-batch size and scheduling time point) to optimize overall performance while ensuring DL convergence accuracy through gradient accumulation. In experiments with both physical DL workloads and trace-driven simulations, even as a preemptionfree policy, SJF-BSBF reduces the average job completion time by 27-33% relative to the state-of-the-art preemptive DL schedulers. Moreover, SJF-BSBF can wisely determine the optimal resource sharing settings, such as the sharing time point and sub-batch size for gradient accumulation, outperforming the aggressive GPU sharing approach (baseline SJF-FFS policy) by up to 17% in large-scale traces.
Yizhou Luo, Qiang Wang 0022, Shaohuai Shi, Jiaxin Lai, Shuhan Qi, Jiajia Zhang 0001, Xuan Wang 0002
IWQoS7
2024 MCDHGN: heterogeneous network-based cancer driver gene prediction and interpretability analysis
abstract
MOTIVATION: Accurately predicting the driver genes of cancer is of great significance for carcinogenesis progress research and cancer treatment. In recent years, more and more deep-learning-based methods have been used for predicting cancer driver genes. However, deep-learning algorithms often have black box properties and cannot interpret the output results. Here, we propose a novel cancer driver gene mining method based on heterogeneous network meta-paths (MCDHGN), which uses meta-path aggregation to enhance the interpretability of predictions. RESULTS: MCDHGN constructs a heterogeneous network by using several types of multi-omics data that are biologically linked to genes. And the differential probabilities of SNV, DNA methylation, and gene expression data between cancerous tissues and normal tissues are extracted as initial features of genes. Nine meta-paths are manually selected, and the representation vectors obtained by aggregating information within and across meta-path nodes are used as new features for subsequent classification and prediction tasks. By comparing with eight homogeneous and heterogeneous network models on two pan-cancer datasets, MCDHGN has better performance on AUC and AUPR values. Additionally, MCDHGN provides interpretability of predicted cancer driver genes through the varying weights of biologically meaningful meta-paths. AVAILABILITY AND IMPLEMENTATION: https://github.com/1160300611/MCDHGN.
Lexiang Wang, Jingli Zhou, Xuan Wang 0002, Yadong Wang 0001, Junyi Li 0004
Bioinform.3
2024 Combining Counterfactual Regret Minimization With Information Gain to Solve Extensive Games With Unknown Environments
abstract
Counterfactual regret minimization (CFR) is an effective algorithm for solving extensive‐form games with imperfect information (IIEGs). However, CFR is only allowed to be applied in known environments, where the transition function of the chance player and the reward function of the terminal node in IIEGs are known. In uncertain situations, such as reinforcement learning (RL) problems, CFR is not applicable. Thus, applying CFR in unknown environments is a significant challenge that can also address some difficulties in the real world. Currently, advanced solutions require more interactions with the environment and are limited by large single‐sampling variances to narrow the gap with the real environment. In this paper, we propose a method that combines CFR with information gain to compute the Nash equilibrium (NE) of IIEGs with unknown environments. We use a curiosity‐driven approach to explore unknown environments and minimize the discrepancy between uncertain and real environments. In addition, by incorporating information into the reward, the average strategy calculated by CFR can be directly implemented as the interaction policy with the environment, thereby improving the exploration efficiency of our method in uncertain environments. Through experiments on standard testbeds such as Kuhn poker and Leduc poker, our method significantly reduces the number of interactions with the environment compared to the different baselines and computes a more accurate approximate NE within the same number of interaction rounds.
Chen Qiu 0003, Xuan Wang 0002, Tianzi Ma, Yaojun Wen, Jiajia Zhang 0001
Int. J. Intell. Syst.2
2024 A fast strategy-solving method for adversarial team games utilizing warm starting
Chen Qiu 0003, Jiajia Zhang 0001, Xuan Wang 0002
Neurocomputing5
2024 CMCL: Cross-Modal Compressive Learning for Resource-Constrained Intelligent IoT Systems
abstract
Compressive Learning (CL) has proven to be highly successful in executing joint signal sampling and inference for intricate vision tasks through resource-limited Internet of Things (IoT) devices. Recent studies have turned their attention towards utilizing the deep neural networks (DNNs) methodology, also known as DeepCL, to enhance performance in unimodal vision tasks. This approach incorporates learnable compressed sensing in a comprehensive, end-to-end manner. Current DeepCL techniques typically employ initial signal reconstruction as the input for subsequent DNNs for inference. However, this practice presents potential risks such as privacy breaches and reduced performance due to information processing inequality. To address these issues, this paper introduces the first cross-modal compressive learning (CMCL) approach that enables image captioning directly on compressed measurements. When compared to previous DeepCL strategies, the proposed CMCL offers significant improvements in computational efficiency and privacy protection. Extensive experiments demonstrate that CMCL performance is nearly on par with leading image captioning methods, showcasing a metric value that is merely 2.75% lower than the uncompressed method when the data is compressed eightfold.
Bin Chen 0011, Yujun Huang, Baoyi An 0002, Yaowei Wang 0001, Xuan Wang 0002
IEEE Internet Things J.6
2024 circ2DGNN: circRNA-Disease Association Prediction via Transformer-Based Graph Neural Network
abstract
Investigating the associations between circRNA and diseases is vital for comprehending the underlying mechanisms of diseases and formulating effective therapies. Computational prediction methods often rely solely on known circRNA-disease data, indirectly incorporating other biomolecules' effects by computing circRNA and disease similarities based on these molecules. However, this approach is limited, as other biomolecules also play significant roles in circRNA-disease interactions. To address this, we construct a comprehensive heterogeneous network incorporating data on human circRNAs, diseases, and other biomolecule interactions to develop a novel computational model, circ2DGNN, which is built upon a heterogeneous graph neural network. circ2DGNN directly takes heterogeneous networks as inputs and obtains the embedded representation of each node for downstream link prediction through graph representation learning. circ2DGNN employs a Transformer-like architecture, which can compute heterogeneous attention score for each edge, and perform message propagation and aggregation, using a residual connection to enhance the representation vector. It uniquely applies the same parameter matrix only to identical meta-relationships, reflecting diverse parameter spaces for different relationship types. After fine-tuning hyperparameters via five-fold cross-validation, evaluation conducted on a test dataset shows circ2DGNN outperforms existing state-of-the-art(SOTA) methods.
Keliang Cen, Zheming Xing, Xuan Wang 0002, Yadong Wang 0001, Junyi Li 0004
IEEE ACM Trans. Comput. Biol. Bioinform.3
2024 A Novel Tree-Based Method for Interpretable Reinforcement Learning
abstract
Deep reinforcement learning (DRL) has garnered remarkable success across various domains, propelled by advancements in deep learning (DL) technologies. However, the opacity of DL presents significant challenges, limiting the application of DRL in critical systems. In response, decision tree (DT)-based methods, known for their transparent decision-making mechanisms, have shown promise in making interpretable policies for decision-making problems. Existing methods often employ differential DTs to model RL policies and discretize them to conventional DTs for higher interpretability. Yet, this method leads to discrepancies between the trained differential DTs and the discretized DTs. To address this issue, we introduce Generative Consistent Trees (GCTs), a novel solution that circumvents the information loss typically associated with the argmax operation in prior research. By implementing a reparameterization technique to approximate the categorical distribution, GCTs ensure the consistencies between trained GCTs and discretized counterparts. Moreover, we have developed an imitation-learning-based framework for interpretable reinforcement learning. This framework is designed to train GCTs by efficiently mimicking expert policies. Our extensive experiments across multiple environments have validated the effectiveness of this approach, highlighting the potential of GCTs in enhancing the interpretability and applicability of DRL.
Shuhan Qi, Xuan Wang 0002, Jiajia Zhang 0001
ACM Trans. Knowl. Discov. Data3
2024 D2CFR: Minimize Counterfactual Regret With Deep Dueling Neural Network
abstract
Counterfactual regret minimization (CFR) is a popular method for finding approximate Nash equilibrium in two-player zero-sum games with imperfect information. Solving large-scale games with CFR needs a combination of abstraction techniques and certain expert knowledge, which constrains its scalability. Recent neural-based CFR methods mitigate the need for abstraction and expert knowledge by training an efficient network to directly obtain counterfactual regret without abstraction. However, these methods only consider estimating regret values for individual actions, neglecting the evaluation of state values, which are significant for decision-making. In this article, we introduce deep dueling CFR (D2CFR), which emphasizes the state value estimation by employing a novel value network with a dueling structure. Moreover, a rectification module based on a time-shifted Monte Carlo simulation is designed to rectify the inaccurate state value estimation. Extensive experimental results are conducted to show that D2CFR converges faster and outperforms comparison methods on test games.
Huale Li, Xuan Wang 0002, Zengyue Guo, Jiajia Zhang 0001, Shuhan Qi
IEEE Trans. Neural Networks Learn. Syst.2
2024 Cascaded Attention: Adaptive and Gated Graph Attention Network for Multiagent Reinforcement Learning
abstract
Modeling the interactive relationships of agents is critical to improving the collaborative capability of a multiagent system. Some methods model these by predefined rules. However, due to the nonstationary problem, the interactive relationship changes over time and cannot be well captured by rules. Other methods adopt a simple mechanism such as an attention network to select the neighbors the current agent should collaborate with. However, in large-scale multiagent systems, collaborative relationships are too complicated to be described by a simple attention network. We propose an adaptive and gated graph attention network (AGGAT), which models the interactive relationships between agents in a cascaded manner. In the AGGAT, we first propose a graph-based hard attention network that roughly filters irrelevant agents. Then, normal soft attention is adopted to decide the importance of each neighbor. Finally, gated attention further refines the collaborative relationship of agents. By using cascaded attention, the collaborative relationship of agents is precisely learned in a coarse-to-fine style. Extensive experiments are conducted on a variety of cooperative tasks. The results indicate that our proposed method outperforms state-of-the-art baselines.
Shuhan Qi, Xinhao Huang, Peixi Peng, Xuzhong Huang, Jiajia Zhang 0001, Xuan Wang 0002
IEEE Trans. Neural Networks Learn. Syst.6
2024 Meta-Path Based Attentional Graph Learning Model for Vulnerability Detection
abstract
In recent years, deep learning (DL)-based methods have been widely used in code vulnerability detection. The DL-based methods typically extract structural information from source code, e.g., code structure graph, and adopt neural networks such as Graph Neural Networks (GNNs) to learn the graph representations. However, these methods fail to consider the heterogeneous relations in the code structure graph, i.e., the heterogeneous relations mean that the different types of edges connect different types of nodes in the graph, which may obstruct the graph representation learning. Besides, these methods are limited in capturing long-range dependencies due to the deep levels in the code structure graph. In this paper, we propose aMeta-path basedAttentionalGraph learning model for code vulNErability deTection, calledMAGNET. MAGNET constructs a multi-granularity meta-path graph for each code snippet, in which the heterogeneous relations are denoted as meta-paths to represent the structural information. A meta-path based hierarchical attentional graph neural network is also proposed to capture the relations between distant nodes in the graph. We evaluate MAGNET on three public datasets and the results show that MAGNET outperforms the best baseline method in terms of F1 score by 6.32%, 21.50%, and 25.40%, respectively. MAGNET also achieves the best performance among all the baseline methods in detecting Top-25 most dangerous Common Weakness Enumerations (CWEs), further demonstrating its effectiveness in vulnerability detection.
Xin-Cheng Wen, Cuiyun Gao 0001, Jiaxin Ye, Yichen Li 0003, Zhihong Tian 0001, Yan Jia 0001, Xuan Wang 0002
IEEE Trans. Software Eng.7
2023 Efficient Cloud Computing Resource Management Strategy Based on Auction Mechanism
Qian Chen 0028, Xuan Wang 0002, Zoe Lin Jiang
APNOMS2
2023 GIFD: A Generative Gradient Inversion Method with Feature Domain Optimization
abstract
Federated Learning (FL) has recently emerged as a promising distributed machine learning framework to preserve clients' privacy, by allowing multiple clients to upload the gradients calculated from their local data to a central server. Recent studies find that the exchanged gradients also take the risk of privacy leakage, e.g., an attacker can invert the shared gradients and recover sensitive data against an FL system by leveraging pre-trained generative adversarial networks (GAN) as prior knowledge. However, performing gradient inversion attacks in the latent space of the GAN model limits their expression ability and generalizability. To tackle these challenges, we propose Gradient Inversion over Feature Domains (GIFD), which disassembles the GAN model and searches the feature domains of the intermediate layers. Instead of optimizing only over the initial latent code, we progressively change the optimized layer, from the initial latent space to intermediate layers closer to the output images. In addition, we design a regularizer to avoid unreal image generation by adding a small l1ball constraint to the searching range. We also extend GIFD to the out-of-distribution (OOD) setting, which weakens the assumption that the training sets of GANs and FL tasks obey the same data distribution. Extensive experiments demonstrate that our method can achieve pixel-level reconstruction and is superior to the existing methods. Notably, GIFD also shows great generalizability under different defense strategy settings and batch sizes.
Hao Fang 0011, Bin Chen 0011, Xuan Wang 0002, Zhi Wang 0001, Shutao Xia
ICCV3
2023 NIEE: Modeling Edge Embeddings for Drug-Disease Association Prediction via Neighborhood Interactions
Yu Jiang 0002, Jingli Zhou, Yulin Wu 0001, Xuan Wang 0002, Junyi Li 0004
ICIC (3)5
2023 HNetGO: protein function prediction via heterogeneous network transformer
abstract
Protein function annotation is one of the most important research topics for revealing the essence of life at molecular level in the post-genome era. Current research shows that integrating multisource data can effectively improve the performance of protein function prediction models. However, the heavy reliance on complex feature engineering and model integration methods limits the development of existing methods. Besides, models based on deep learning only use labeled data in a certain dataset to extract sequence features, thus ignoring a large amount of existing unlabeled sequence data. Here, we propose an end-to-end protein function annotation model named HNetGO, which innovatively uses heterogeneous network to integrate protein sequence similarity and protein-protein interaction network information and combines the pretraining model to extract the semantic features of the protein sequence. In addition, we design an attention-based graph neural network model, which can effectively extract node-level features from heterogeneous networks and predict protein function by measuring the similarity between protein nodes and gene ontology term nodes. Comparative experiments on the human dataset show that HNetGO achieves state-of-the-art performance on cellular component and molecular function branches.
Xiaoshuai Zhang, Huannan Guo, Xuan Wang 0002, Kaitao Wu, Shizheng Qiu, Bo Liu 0023, Yadong Wang 0001, Yang Hu 0008, Junyi Li 0004
Briefings Bioinform.4
2023 Struct2GO: protein function prediction based on graph pooling algorithm and AlphaFold2 structure information
abstract
MOTIVATION: In recent years, there has been a breakthrough in protein structure prediction, and the AlphaFold2 model of the DeepMind team has improved the accuracy of protein structure prediction to the atomic level. Currently, deep learning-based protein function prediction models usually extract features from protein sequences and combine them with protein-protein interaction networks to achieve good results. However, for newly sequenced proteins that are not in the protein-protein interaction network, such models cannot make effective predictions. To address this, this article proposes the Struct2GO model, which combines protein structure and sequence data to enhance the precision of protein function prediction and the generality of the model. RESULTS: We obtain amino acid residue embeddings in protein structure through graph representation learning, utilize the graph pooling algorithm based on a self-attention mechanism to obtain the whole graph structure features, and fuse them with sequence features obtained from the protein language model. The results demonstrate that compared with the traditional protein sequence-based function prediction model, the Struct2GO model achieves better results. AVAILABILITY AND IMPLEMENTATION: The data underlying this article are available at https://github.com/lyjps/Struct2GO.
Peishun Jiao, Xuan Wang 0002, Bo Liu 0023, Yadong Wang 0001, Junyi Li 0004
Bioinform.3
2023 SLGNN: synthetic lethality prediction in human cancers based on factor-aware knowledge graph neural network
abstract
MOTIVATION: Synthetic lethality (SL) is a form of genetic interaction that can selectively kill cancer cells without damaging normal cells. Exploiting this mechanism is gaining popularity in the field of targeted cancer therapy and anticancer drug development. Due to the limitations of identifying SL interactions from laboratory experiments, an increasing number of research groups are devising computational prediction methods to guide the discovery of potential SL pairs. Although existing methods have attempted to capture the underlying mechanisms of SL interactions, methods that have a deeper understanding of and attempt to explain SL mechanisms still need to be developed. RESULTS: In this work, we propose a novel SL prediction method, SLGNN. This method is based on the following assumption: SL interactions are caused by different molecular events or biological processes, which we define as SL-related factors that lead to SL interactions. SLGNN, apart from identifying SL interaction pairs, also models the preferences of genes for different SL-related factors, making the results more interpretable for biologists and clinicians. SLGNN consists of three steps: first, we model the combinations of relationships in the gene-related knowledge graph as the SL-related factors. Next, we derive initial embeddings of genes through an explicit message aggregation process of the knowledge graph. Finally, we derive the final gene embeddings through an SL graph, constructed using known SL gene pairs, utilizing factor-based message aggregation. At this stage, a supervised end-to-end training model is used for SL interaction prediction. Based on experimental results, the proposed SLGNN model outperforms all current state-of-the-art SL prediction methods and provides better interpretability. AVAILABILITY AND IMPLEMENTATION: SLGNN is freely available at https://github.com/zy972014452/SLGNN.
Yuhuan Zhou, Yang Liu 0039, Xuan Wang 0002, Junyi Li 0004
Bioinform.4
2023 What is the limitation of multimodal LLMs? A deeper look into multimodal LLMs through prompt probing
Shuhan Qi, Zhengying Cao, Jun Rao, Lei Wang 0203, Jing Xiao 0006, Xuan Wang 0002
Inf. Process. Manag.6
2023 Kdb-D2CFR: Solving Multiplayer imperfect-information games with knowledge distillation-based DeepCFR
Huale Li, Zengyue Guo, Yang Liu 0039, Xuan Wang 0002, Shuhan Qi, Jiajia Zhang 0001, Jing Xiao 0006
Knowl. Based Syst.4
2023 Breaking the traditional: a survey of algorithmic mechanism design applied to economic and complex environments
Qian Chen 0028, Xuan Wang 0002, Zoe Lin Jiang, Yulin Wu 0001, Huale Li, Xiaozhen Sun
Neural Comput. Appl.2
2023 Learning to Generate Tips from Song Reviews
Jingya Zang, Cuiyun Gao 0001, Yupan Chen, Ruifeng Xu 0001, Lanjun Zhou, Xuan Wang 0002
Neural Networks6
2023 Adversarial Examples Generation for Deep Product Quantization Networks on Image Retrieval
abstract
Deep product quantization networks (DPQNs) have been successfully used in image retrieval tasks, due to their powerful feature extraction ability and high efficiency of encoding high-dimensional visual features. Recent studies show that deep neural networks (DNNs) are vulnerable to input with small and maliciously designed perturbations (a.k.a., adversarial examples) for classification. However, little effort has been devoted to investigating how adversarial examples affect DPQNs, which raises the potential safety hazard when deploying DPQNs in a commercial search engine. To this end, we propose an adversarial example generation framework by generating adversarial query images for DPQN-based retrieval systems. Unlike the adversarial generation for the classic image classification task that heavily relies on ground-truth labels, we alternatively perturb the probability distribution of centroids assignments for a clean query, then we can induce effective non-targeted attacks on DPQNs in white-box and black-box settings. Moreover, we further extend the non-targeted attack to a targeted attack by a novel sample space averaging scheme ([Formula: see text]AS), whose theoretical guarantee is also obtained. Extensive experiments show that our methods can create adversarial examples to successfully mislead the target DPQNs. Besides, we found that our methods both significantly degrade the retrieval performance under a wide variety of experimental settings. The source code is available at https://github.com/Kira0096/PQAG.
Bin Chen 0011, Tao Dai 0001, Jiawang Bai, Yong Jiang 0001, Shutao Xia, Xuan Wang 0002
IEEE Trans. Pattern Anal. Mach. Intell.7
2023 A blockchain-based privacy-preserving advertising attribution architecture: Requirements, design, and a prototype implementation
abstract
Abstract In the era of digital marketing, advertisements have become an indispensable part. One of the central challenges is advertising attribution which explains the amount of contribution every publisher has with the conversions. However, through observation, we have found that current advertising platforms attribution, advertisers attribution, or third‐party platforms attribution all have the problems of trust, data leakage, and data forgery. To fill the gap, our work's main contribution is combining blockchain with advertising attribution to propose an architecture for improving the privacy‐preserving degree and amount. In the proposed architecture, publishers, and advertisers can store real‐time data on a blockchain. The attribution results are credible because blockchain is decentralized, tamper‐proof, and traceable. We combine privacy set intersection and zero‐knowledge proof technology to increase the privacy of flowing data. In addition, we describe a preliminary prototype in which publishers, advertisers, and advertising platforms can get the corresponding attribution details. To show its effectiveness, we analyze it from different perspectives, including communication cost, attribution accuracy, and time cost. The results show that our communication cost has significantly reduced compared to the recent studies.
Yang Liu 0039, Liangjie Lin, Weizhe Zhang, Xuan Wang 0002, Mehdi Gheisari, Hamid Esmaeili Najafabadi
Softw. Pract. Exp.5
2023 The Design of Fast and Lightweight Resemblance Detection for Efficient Post-Deduplication Delta Compression
abstract
Post-deduplication delta compression is a data reduction technique that calculates and stores the differences of very similar but non-duplicate chunks in storage systems, which is able to achieve a very high compression ratio. However, the low throughput of widely used resemblance detection approaches (e.g., N-Transform) usually becomes the bottleneck of delta compression systems due to introducing high computational overhead. Generally, this overhead mainly consists of two parts: ① calculating the rolling hash byte by byte across data chunks and ② applying multiple transforms on all of the calculated rolling hash values. In this article, we propose Odess, a fast and lightweight resemblance detection approach, that greatly reduces the computational overhead for resemblance detection while achieving high detection accuracy and a high compression ratio. Odess first utilizes a novel Subwindow-based Parallel Rolling (SWPR) hash method using Single Instruction Multiple Data [ 1 ] (SIMD) to accelerate calculation of rolling hashes (corresponding to the first part of the overhead). Odess then uses a novel Content-Defined Sampling method to generate a much smaller proxy hash set from the whole rolling hash set and quickly applies transforms on this small hash set for resemblance detection (corresponding to the second part of the overhead). Evaluation results show that during the stage of resemblance detection, the Odess approach is ∼31.4× and ∼7.9× faster than the state-of-the-art N-Transform and Finesse (a recent variant of N-Transform [ 39 ]), respectively. When considering an end-to-end data reduction storage system, the Odess-based system’s throughput is about 3.20× and 1.41× higher than the N-Transform- and Finesse-based systems’ throughput, respectively, while maintaining the high compression ratio of N-Transform and achieving ∼1.22× higher compression ratio over Finesse.
Wen Xia, Lifeng Pu, Xiangyu Zou, Philip Shilane, Haijun Zhang 0002, Xuan Wang 0002
ACM Trans. Storage7
2023 Listening to Users' Voice: Automatic Summarization of Helpful App Reviews
abstract
App reviews are crowdsourcing knowledge of user experience with the apps, providing valuable information for app release planning, such as major bugs to fix and important features to add. There exist prior explorations on app review mining for release planning; however, most of the studies strongly rely on predefined classes or manually annotated reviews. Also, the new review characteristic, i.e., the number of users who rated the review as helpful, which can help capture important reviews, has not been considered previously. In the article, we propose a novel framework, named SOLAR, aiming at accurately summarizing helpful user reviews to developers. The framework mainly contains three modules: the review helpfulness prediction module, topic-sentiment modeling module, and multifactor ranking module. The review helpfulness prediction module assesses the helpfulness of reviews, i.e., whether the review is useful for developers. The topic-sentiment modeling module groups the topics of the helpful reviews and also predicts the associated sentiment, and the multifactor ranking module aims at prioritizing semantically representative reviews for each topic as the review summary. Experiments on five popular apps indicate that SOLAR is effective for review summarization and promising for facilitating app release planning.
Cuiyun Gao 0001, Shuhan Qi, Yang Liu 0039, Xuan Wang 0002, Zibin Zheng, Qing Liao 0001
IEEE Trans. Reliab.5
2022 Underwater Small Target Detection Based on Deformable Convolutional Pyramid
abstract
Due to the problem of severe deformation, occlusion, diversified scenarios, general object detection methods cannot achieve satisfactory results in underwater object detection tasks. In this paper, we propose a two-stage Underwater Small Target Detection (USTD) network. In the proposed USTD, the Deformable Convolutional Pyramid(DCP) is proposed to deal with the problems of deformation, occlusion, and various object sizes effectively. Besides, we also propose a strategy of domain generalization based on curriculum learning to improve generalization in multi-domain environments, which is named as Phased Learning. Afterward, we construct an underwater target detection set (UTDS) to evaluate the accuracy of our method in underwater target detection tasks. Our method shows superior detection performance in experiments and reaches state-of-the-art for underwater target detection. Finally, in the 2020 China Underwater Robot Professional Contest (URPC), our method reached third place in terms of accuracy.
Shuhan Qi, Jianjun Du, Mingyan Wu, Linlin Tang, Tao Qian 0003, Xuan Wang 0002
ICASSP7
2022 Building a High-performance Fine-grained Deduplication Framework for Backup Storage with High Deduplication Ratio
Xiangyu Zou, Wen Xia, Philip Shilane, Haijun Zhang 0002, Xuan Wang 0002
USENIX ATC5
2022 RLCFR: Minimize counterfactual regret by deep reinforcement learning
Huale Li, Xuan Wang 0002, Fengwei Jia, Yulin Wu 0001, Jiajia Zhang 0001, Shuhan Qi
Expert Syst. Appl.2
2022 Parallel learner: A practical deep reinforcement learning framework for multi-scenario games
Xiaohan Hou, Zhenyang Guo, Xuan Wang 0002, Tao Qian 0003, Jiajia Zhang 0001, Shuhan Qi, Jing Xiao 0006
Knowl. Based Syst.3
2022 Practical protection against video data leakage via universal adversarial head
Jiawang Bai, Bin Chen 0011, Kuofeng Gao, Xuan Wang 0002, Shutao Xia
Pattern Recognit.4
2022 An Integrated Multi-Task Model for Fake News Detection
abstract
Fake news detection attracts many researchers’ attention due to the negative impacts on the society. Most existing fake news detection approaches mainly focus on semantic analysis of news’ contents. However, the detection performance will dramatically decrease when the content of news is short. In this paper, we propose a novelfake news detection multi-task learning (FDML)model based on the following observations: 1) some certain topics have higher percentages of fake news; and 2) some certain news authors have higher intentions to publish fake news. FDML model investigates the impact of topic labels for the fake news and introduce contextual information of news at the same time to boost the detection performance on the short fake news. Specifically, the FDML model consists of representation learning and multi-task learning parts to train the fake news detection task and the news topic classification task, simultaneously. As far as we know, this is the first fake news detection work that integrates the above two tasks. The experiment results show that the FDML model outperforms state-of-the-art methods on real-world fake news dataset.
Qing Liao 0001, Heyan Chai 0001, Xiang Zhang 0008, Xuan Wang 0002, Wen Xia, Ye Ding 0002
IEEE Trans. Knowl. Data Eng.5
2022 From Hyper-dimensional Structures to Linear Structures: Maintaining Deduplicated Data's Locality
abstract
Data deduplication is widely used to reduce the size of backup workloads, but it has the known disadvantage of causing poor data locality, also referred to as the fragmentation problem. This results from the gap between the hyper-dimensional structure of deduplicated data and the sequential nature of many storage devices, and this leads to poor restore and garbage collection (GC) performance. Current research has considered writing duplicates to maintain locality (e.g., rewriting) or caching data in memory or SSD, but fragmentation continues to lower restore and GC performance. Investigating the locality issue, we design a method to flatten the hyper-dimensional structured deduplicated data to a one-dimensional format, which is based on classification of each chunk’s lifecycle, and this creates our proposed data layout. Furthermore, we present a novel management-friendly deduplication framework, called MFDedup, that applies our data layout and maintains locality as much as possible. Specifically, we use two key techniques in MFDedup: Neighbor-duplicate-focus indexing (NDF) and Across-version-aware Reorganization scheme (AVAR). NDF performs duplicate detection against a previous backup, then AVAR rearranges chunks with an offline and iterative algorithm into a compact, sequential layout, which nearly eliminates random I/O during file restores after deduplication. Evaluation results with five backup datasets demonstrate that, compared with state-of-the-art techniques, MFDedup achieves deduplication ratios that are 1.12× to 2.19× higher and restore throughputs that are 1.92× to 10.02× faster due to the improved data layout. While the rearranging stage introduces overheads, it is more than offset by a nearly-zero overhead GC process. Moreover, the NDF index only requires indices for two backup versions, while the traditional index grows with the number of versions retained.
Xiangyu Zou, Jingsong Yuan, Philip Shilane, Wen Xia, Haijun Zhang 0002, Xuan Wang 0002
ACM Trans. Storage6
2022 NetSync: A Network Adaptive and Deduplication-Inspired Delta Synchronization Approach for Cloud Storage Services
abstract
Delta sync (synchronization) is a key bandwidth-saving technique for cloud storage services. The representative delta sync utility,rsync, matches data chunks by sliding a search window byte-by-byte to maximize the redundancy detection for bandwidth efficiency. However, it is difficult for this process to cater to the forthcoming high-bandwidth cloud storage services which require lightweight delta sync that can well support large files. Moreover,rsyncemploys invariant chunking and compression methods during the sync process, making it unable to cater to services from various network environments which require the sync approach to perform well under different network conditions. Inspired by the Content-Defined Chunking (CDC) technique used in data deduplication, we propose NetSync, a network adaptive and CDC-based lightweight delta sync approach with less computing and protocol (metadata) overheads than the state-of-the-art delta sync approaches. Besides, NetSync can choose appropriate compressing and chunking strategies for different network conditions. The key idea of NetSync is (1) to simplify the process of chunk matching by proposing a fast weak hash called FastFP that is piggybacked on the rolling hashes from CDC, and redesigning the delta sync protocol by exploiting deduplication locality and weak/strong hash properties; (2) to minimize the sync time by adaptively choosing chunking parameters and compression methods according to the current network conditions. Our evaluation results driven by both benchmark and real-world datasets suggest NetSync performs$2\times$2×–$10\times$10×faster and supports$30\%$30%–$80\%$80%more clients than the state-of-the-artrsync-based WebR2sync+ and deduplication-based approach.
Wen Xia, Can Wei, Zhenhua Li 0001, Xuan Wang 0002, Xiangyu Zou
IEEE Trans. Parallel Distributed Syst.4
2021 Student Can Also be a Good Teacher: Extracting Knowledge from Vision-and-Language Model for Cross-Modal Retrieval
abstract
Astounding results from transformer models with Vision-and Language Pretraining (VLP) on joint vision-and-language downstream tasks have intrigued the multi-modal community. On the one hand, these models are usually so huge that make us more difficult to fine-tune and serve real-time online applications. On the other hand, the compression of the original transformer block will ignore the difference in information between modalities, which leads to the sharp decline of retrieval accuracy.
Jun Rao, Tao Qian 0003, Shuhan Qi, Yulin Wu 0001, Qing Liao 0001, Xuan Wang 0002
CIKM6
2021 The Dilemma between Deduplication and Locality: Can Both be Achieved?
Xiangyu Zou, Jingsong Yuan, Philip Shilane, Wen Xia, Haijun Zhang 0002, Xuan Wang 0002
FAST6
2021 FedSP: Federated Speaker Verification with Personal Privacy Preservation
Yangqian Wang, Yuanfeng Song, Di Jiang 0004, Ye Ding 0002, Xuan Wang 0002, Yang Liu 0039, Qing Liao 0001
ICA3PP (3)5
2021 Odess: Speeding up Resemblance Detection for Redundancy Elimination by Fast Content-Defined Sampling
abstract
Multiple data reduction techniques have been investigated to lower storage costs for a wide variety of customers. In this work, we focus on similarity-based delta compression, which calculates and stores the difference of very similar, but non-duplicate, chunks in storage systems. Delta compression is often implemented along with deduplication and has been shown to achieve a much higher compression ratio. Currently, the N-Transform method is the most popular and widely-used approach to generate features for data content (e.g. chunks) to detect similar candidates (and then apply delta compression). For delta compression systems, though, the throughput of N-Transform is often the bottleneck. Finesse is a high throughput variant of N-Transform, but it suffers from lower detection accuracy and compression ratio. The computation overhead of N-Transform consists of two parts: calculating the rolling hash across data and applying time-consuming transforms on each hash. In this work, we propose Odess, a fast resemblance detection approach, that uses a novel Content-Defined Sampling method to generate a much smaller proxy hash set and then applies transforms on this small hash set. This reduces the calculations in the transform step from being the bottleneck. Meanwhile, Odess also leverages the faster Gear hash to generate rolling hashes. Thus, Odess greatly reduces the computational overhead for resemblance detection while achieving high detection accuracy and high compression ratio. Our evaluation results show that Odess is ~ 5.4× (Finesse) and ~ 26.9× (N-Transform) faster (on average) at generating features for resemblance detection. When considering an end-to-end data reduction storage system, Odess increases throughput by ~ 1.36× (Finesse) and ~ 2.76× (N-Transform) while maintaining the compression ratio of N-Transform and increasing the compression ratio ~ 1.22× over Finesse.
Xiangyu Zou, Wen Xia, Philip Shilane, Haoliang Tan, Haijun Zhang 0002, Xuan Wang 0002
ICDE7
2021 Fine-Grained Unbalanced Interaction Network for Visual Question Answering
Xinxin Liao, Mingyan Wu, Heyan Chai 0001, Shuhan Qi, Xuan Wang 0002, Qing Liao 0001
KSEM5
2021 Reconstruct Anomaly to Normal: Adversarially Learned and Latent Vector-Constrained Autoencoder for Time-Series Anomaly Detection
Chunkai Zhang, Wei Zuo, Shaocong Li, Xuan Wang 0002, Peiyi Han, Chuanyi Liu
PRICAI (2)4
2021 WRGPruner: A new model pruning solution for tiny salient object detection
Fengwei Jia, Xuan Wang 0002, Jian Guan 0001, Huale Li, Chen Qiu 0003, Shuhan Qi
Image Vis. Comput.2
2021 ARank: Toward specific model pruning via advantage rank for multiple salient objects detection
Fengwei Jia, Xuan Wang 0002, Jian Guan 0001, Huale Li, Chen Qiu 0003, Shuhan Qi
Image Vis. Comput.2
2021 Scalable sub-game solving for imperfect-information games
Huale Li, Xuan Wang 0002, Kunchi Li, Fengwei Jia, Yulin Wu 0001, Jiajia Zhang 0001, Shuhan Qi
Knowl. Based Syst.2
2021 Autoencoder-based self-supervised hashing for cross-modal retrieval
Xuan Wang 0002, Jiajia Zhang 0001, Chengkai Huang, Shuhan Qi
Multim. Tools Appl.2
2021 Efficient Server-Aided Secure Two-Party Computation in Heterogeneous Mobile Cloud Computing
abstract
With the ubiquity of mobile devices and rapid development of cloud computing, mobile cloud computing (MCC) has been considered as an essential computation setting to support complicated, scalable and flexible mobile applications by overcoming the physical limitations of mobile devices with the aid of cloud. In the MCC setting, since many mobile applications (e.g., map apps) interacting with cloud server and application server need to perform computation with the private data of users, it is important to realize secure computation for MCC. In this article, we propose an efficient server-aided secure two-party computation (2PC) protocol for MCC. This is the first work that considers collusion between a malicious garbled circuit evaluator and a semi-honest server while ensuring privacy and correctness. Also, it can guarantee fairness when collusion does not exist. The security analysis shows that our protocol can securely compute any function f(x, y) against different types of adversaries in the malicious model. Also, the experimental performance analysis shows that this work outperforms the previous works for at least 10 times with the same security level.
Yulin Wu 0001, Xuan Wang 0002, Willy Susilo, Guomin Yang, Zoe Lin Jiang, Qian Chen 0028, Peng Xu 0003
IEEE Trans. Dependable Secur. Comput.2
2021 Constraint-Objective Cooperative Coevolution for Large-scale Constrained Optimization
abstract
Large-scale optimization problems and constrained optimization problems have attracted considerable attention in the swarm and evolutionary intelligence communities and exemplify two common features of real problems, i.e., a large scale and constraint limitations. However, only a little work on solving large-scale continuous constrained optimization problems exists. Moreover, the types of benchmarks proposed for large-scale continuous constrained optimization algorithms are not comprehensive at present. In this article, first, a constraint-objective cooperative coevolution (COCC) framework is proposed for large-scale continuous constrained optimization problems, which is based on the dual nature of the objective and constraint functions: modular and imbalanced components. The COCC framework allocates the computing resources to different components according to the impact of objective values and constraint violations. Second, a benchmark for large-scale continuous constrained optimization is presented, which takes into account the modular nature, as well as both imbalanced and overlapping characteristics of components. Finally, three different evolutionary algorithms are embedded into the COCC framework for experiments, and the experimental results show that COCC performs competitively.
Peilan Xu, Wenjian Luo, Xin Lin 0004, Jiajia Zhang 0001, Yingying Qiao, Xuan Wang 0002
ACM Trans. Evol. Learn. Optim.6
2020 Enhanced Gaze Following via Object Detection and Human Pose Estimation
Jian Guan 0001, Liming Yin, Shuhan Qi, Xuan Wang 0002, Qing Liao 0001
MMM (2)5
2020 A mix-supervised unified framework for salient object detection
Fengwei Jia, Jian Guan 0001, Shuhan Qi, Huale Li, Xuan Wang 0002
Appl. Intell.5
2020 Bi-Connect Net for salient object detection
Fengwei Jia, Xuan Wang 0002, Jian Guan 0001, Qing Liao 0001, Jiajia Zhang 0001, Huale Li, Shuhan Qi
Neurocomputing2
2020 Explore instance similarity: An instance correlation based hashing method for multi-label cross-model retrieval
Chengkai Huang, Jiajia Zhang 0001, Qing Liao 0001, Xuan Wang 0002, Zoe Lin Jiang, Shuhan Qi
Inf. Process. Manag.5
2020 Efficient two-party privacy-preserving collaborative k-means clustering protocol supporting both storage and computation outsourcing
Zoe Lin Jiang, Yabin Jin, Jiazhuo Lv, Yulin Wu 0001, Zechao Liu, Siu-Ming Yiu, Xuan Wang 0002
Inf. Sci.9
2020 Performance Optimization for Relative-Error-Bounded Lossy Compression on Scientific Data
abstract
Scientific simulations in high-performance computing (HPC) environments generate vast volume of data, which may cause a severe I/O bottleneck at runtime and a huge burden on storage space for postanalysis. Unlike traditional data reduction schemes such as deduplication or lossless compression, not only can error-controlled lossy compression significantly reduce the data size but it also holds the promise to satisfy user demand on error control. Pointwise relative error bounds (i.e., compression errors depends on the data values) are widely used by many scientific applications with lossy compression since error control can adapt to the error bound in the dataset automatically. Pointwise relative-error-bounded compression is complicated and time consuming. We develop efficient precomputation-based mechanisms based on the SZ lossy compression framework. Our mechanisms can avoid costly logarithmic transformation and identify quantization factor values via a fast table lookup, greatly accelerating the relative-error-bounded compression with excellent compression ratios. In addition, we reduce traversing operations for Huffman decoding, significantly accelerating the decompression process in SZ. Experiments with eight well-known real-world scientific simulation datasets show that our solution can improve the compression and decompression rates (i.e., the speed) by about 40 and 80 p, respectively, in most of cases, making our designed lossy compression strategy the best-in-class solution in most cases.
Xiangyu Zou, Tao Lu 0014, Wen Xia, Xuan Wang 0002, Weizhe Zhang, Haijun Zhang 0002, Sheng Di, Dingwen Tao, Franck Cappello
IEEE Trans. Parallel Distributed Syst.4
2020 DP-FL: a novel differentially private federated learning framework for the unbalanced data
Xixi Huang, Ye Ding 0002, Zoe Lin Jiang, Shuhan Qi, Xuan Wang 0002, Qing Liao 0001
World Wide Web5
2019 RGB-D tracker under Hierarchical structure
abstract
How to track the target robustly is a challenging task in the field of computer vision. Occlusion as one of the most difficult problems, occurs due to the information lost when three-dimensional subjects are projected in two-dimensional interface, therefore, the 2D or 3D tracking algorithms which adopted depth information that expects to rely on three-dimensional special structure to resolve these problems and made somewhat progress. The 2D tracking algorithm is not efficient in fully using depth information, and the 3D tracking method is not robust because of the lack of mature 3D feature extraction method, which fairly restricts the actual tracking effect. Responding to above questions, we propose an adoption of adaptive quantified depth information, establish an adaptive hierarchical structure according to various scenarios. Hierarchical structure can filter the foreground and background information to reduce the interference in tracking, at the same time simplify the use of the depth information. Combined with kernel correlation filter tracking method, we design the algorithm using 2D apparent model under the spatial structures, which is efficient to deal with the problems of occlusion and the change of target scale, and prove its effectiveness on Princeton Tracking Dataset.
Xuan Wang 0002, Zoe Lin Jiang, Shuhan Qi, Qian Chen 0028
CIFEr2
2019 A New Robust and Reversible Watermarking Technique Based on Erasure Code
Heyan Chai 0001, Shuqiang Yang, Zoe Lin Jiang, Xuan Wang 0002, Hengyu Luo
ICA3PP (1)4
2019 An Adversarial Attack Based on Multi-objective Optimization in the Black-Box Scenario: MOEA-APGA II
Chunkai Zhang, Yepeng Deng, Xuan Wang 0002, Chuanyi Liu
ICICS4
2019 Solving Six-Player Games via Online Situation Estimation
abstract
While the artificial intelligence theory for solving the perfect-information games has been well developed in recent years, great challenges are still posed in dealing with the imperfect-information game due to the huge state space and hidden information involved in it. In this paper, we design an online strategy solving framework for six-player no-limit Texas hold'em poker. Based on hand isomorphism and hand strength evalution, the framework provides an efficient situation estimation method for six-player poker. Such method could greatly reduce the the state space in six-player poker as well as effectively evaluate the current hands. The poker agent based on our method won the third place in the 2018 AAAI-ACPC.
Huale Li, Xuan Wang 0002, Shuhan Qi, Yang Liu 0039, Fengwei Jia, Jiajia Zhang 0001
ICTAI2
2019 Accelerating Relative-error Bounded Lossy Compression for HPC datasets with Precomputation-Based Mechanisms
abstract
Scientific simulations in high-performance computing (HPC) environments are producing vast volume of data, which may cause a severe I/O bottleneck at runtime and a huge burden on storage space for post-analysis. Unlike the traditional data reduction schemes (such as deduplication or lossless compression), not only can error-controlled lossy compression significantly reduce the data size but it can also hold the promise to satisfy user demand on error control. Point-wise relative error bounds (i.e., compression errors depends on the data values) are widely used by many scientific applications in the lossy compression, since error control can adapt to the precision in the dataset automatically. Point-wise relative error bounded compression is complicated and time consuming. In this work, we develop efficient precomputation-based mechanisms in the SZ lossy compression framework. Our mechanisms can avoid costly logarithmic transformation and identify quantization factor values via a fast table lookup, greatly accelerating the relative-error bounded compression with excellent compression ratios. In addition, our mechanisms also help reduce traversing operations for Huffman decoding, and thus significantly accelerate the decompression process in SZ. Experiments with four well-known real-world scientific simulation datasets show that our solution can improve the compression rate by about 30% and decompression rate by about 70% in most of cases, making our designed lossy compression strategy the best choice in class in most cases.
Xiangyu Zou, Tao Lu 0014, Wen Xia, Xuan Wang 0002, Weizhe Zhang, Sheng Di, Dingwen Tao, Franck Cappello
MSST4
2019 Multi-depth Graph Convolutional Networks for Fake News Detection
Guoyong Hu, Ye Ding 0002, Shuhan Qi, Xuan Wang 0002, Qing Liao 0001
NLPCC (1)4
2019 SSHTDNS: A Secure, Scalable and High-Throughput Domain Name System via Blockchain Technique
Zhentian Xiong, Zoe Lin Jiang, Shuqiang Yang, Xuan Wang 0002
NSS4
2019 Bi-directional Features Reuse Network for Salient Object Detection
Fengwei Jia, Xuan Wang 0002, Jian Guan 0001, Shuhan Qi, Qing Liao 0001, Huale Li
PRICAI (3)2
2019 A Lattice-Based Anonymous Distributed E-Cash from Bitcoin
Zeming Lu, Zoe Lin Jiang, Yulin Wu 0001, Xuan Wang 0002, Yantao Zhong
ProvSec4
2019 Multi-task deep convolutional neural network for cancer diagnosis
Qing Liao 0001, Ye Ding 0002, Zoe Lin Jiang, Xuan Wang 0002, Chunkai Zhang, Qian Zhang 0001
Neurocomputing4
2019 Fast Eclat Algorithms Based on Minwise Hashing for Large Scale Transactions
abstract
The Eclat algorithm is one of the most widely used frequent itemset mining methods. In the normal Eclat algorithm and its variants, it is inefficient to calculate the intersection size of itemsets by sequentially comparing elements, especially for large scale transactions. In this paper, we propose the fast Eclat algorithms that can quickly calculate the intersection size of multiple itemsets by using minwise hashing and the estimators. Minwise hashing is used to calculate the Jaccard similarity coefficient by mapping the elements of the sets to those of smaller sets. Two estimators are used to estimate the intersection size of itemsets based on the Jaccard similarity coefficient. Due to the “imperfect” hash function, minwise hashing may obtain a biased Jaccard similarity, which results in error between the real value and the estimated value of the intersection size. Thus, we proposed the HashEclat which uses the maximum of |A| and |B| to represent the union size |A U B|, and proposed the Sim-Eclat which uses the minimum of |A| and |B| to represent the intersection size |A fl B|. Furthermore, we use a boundary error E for better performance as follows: if E is large, the intersection size is determined by a traditional method, and the result is more accurate but takes longer to compute; otherwise, it will be the opposite. Both the theoretical analysis and experimental results show that the proposed algorithms can obtain almost all frequent itemsets with higher speed and less memory usage than other algorithms.
Chunkai Zhang, Panbo Tian, Zoe Lin Jiang, Lin Yao 0004, Xuan Wang 0002
IEEE Internet Things J.6
2019 Graph-based supervised discrete image hashing
Jian Guan 0001, Xuan Wang 0002, Hainan Zhao, Jiajia Zhang 0001, Zechao Liu, Shuhan Qi
J. Vis. Commun. Image Represent.4
2019 Large scale product search with spatial quantization and deep ranking
Shuhan Qi, Zawlin Kyaw, Xuan Wang 0002, Zoe Lin Jiang, Jian Guan 0001
Multim. Tools Appl.3
2018 Variational Autoregressive Decoder for Neural Response Generation
abstract
Combining the virtues of probability graphic models and neural networks, Conditional Variational Auto-encoder (CVAE) has shown promising performance in many applications such as response generation.However, existing CVAE-based models often generate responses from a single latent variable which may not be sufficient to model high variability in responses.To solve this problem, we propose a novel model that sequentially introduces a series of latent variables to condition the generation of each word in the response sequence.In addition, the approximate posteriors of these latent variables are augmented with a backward Recurrent Neural Network (RNN), which allows the latent variables to capture long-term dependencies of future tokens in generation.To facilitate training, we supplement our model with an auxiliary objective that predicts the subsequent bag of words.Empirical experiments conducted on the OpenSubtitle and Reddit datasets show that the proposed model leads to significant improvements on both relevance and diversity over state-of-the-art baselines.
Jiachen Du, Wenjie Li 0002, Yulan He 0001, Ruifeng Xu 0001, Lidong Bing, Xuan Wang 0002
EMNLP6
2018 Efficient Two-Party Privacy Preserving Collaborative k-means Clustering Protocol Supporting both Storage and Computation Outsourcing
Zoe Lin Jiang, Yabin Jin, Jiazhuo Lv, Yulin Wu 0001, Yating Yu, Xuan Wang 0002, Siu-Ming Yiu
ICA3PP (4)7
2018 Outsourced Privacy Preserving SVM with Multiple Keys
Wenli Sun, Zoe Lin Jiang, Jun Zhang 0049, Siu-Ming Yiu, Yulin Wu 0001, Hainan Zhao, Xuan Wang 0002, Peng Zhang 0029
ICA3PP (4)7
2018 Towards Secure Cloud Data Similarity Retrieval: Privacy Preserving Near-Duplicate Image Data Detection
Yulin Wu 0001, Xuan Wang 0002, Zoe Lin Jiang, Xuan Li 0007, Jin Li 0002, Siu-Ming Yiu, Zechao Liu, Hainan Zhao, Chunkai Zhang
ICA3PP (4)2
2018 PPLDEM: A Fast Anomaly Detection Algorithm with Privacy Preserving
Ao Yin, Chunkai Zhang, Zoe Lin Jiang, Yulin Wu 0001, Keli Zhang, Xuan Wang 0002
ICA3PP (4)7
2018 Pixel Meets Region: A Pratical Framework for Salient Object Detection
abstract
Due to the development of deep learning and Fully Convolutional Neural Network (FCN), the research on salient object detection has made great progress in recent years. However, such FCN based models are always affected by the scale-space problem, which reduces the saliency detection accuracy and leads to a blurred object boundary. In this paper, we propose a novel saliency object detection method, which predicts the saliency by incorporating both pixel-level and region-level predictions. First, in order to alleviate the scale-space problem, modified dilated convolution layers and short connections are integrated into the FCN model, and the pixel-level saliency maps is generated by a pixel-wise salient classifier. Then, we employ a superpixel based manifold learning algorithm to obtain a better boundary of salient object, by which a region-level saliency map with clearer object boundary is generated. At last, a simple fusion method is utilized to fuse the two saliency maps into a unified saliency map, followed by a DenseCRF post-refinement module to further optimize the final results. Experiments are conducted on two benchmark datasets to demonstrate the effectiveness of our method.
Xuan Wang 0002, Shuhan Qi, Jian Guan 0001, Fengwei Jia, Lin Yao 0004
ICME2
2018 Balance the Loss: Improving Deep Hash via Loss Weighting and Semantic Preserving
abstract
Learning to hash is widely used in approximate nearest-neighbor (ANN) search. However, traditional hash learning methods, which split the hashing into two parts: feature extraction and hash function learning, usually result in a low retrieval accuracy. Although existing deep learning based hashing methods can improve hashing quality by coupling feature learning and hash encoding, they are always affected by the positive-negative sample imbalance problem. It often deteriorates the performance of the generated hash code. In this paper, we propose an end-to-end deep hashing framework, in which a weighted pairwise loss function is employed to alleviate sample imbalance problem. The loss generated by the positive pairs and negative pairs are given different weights automatically. Moreover, we integrate a classification network into the hashing framework, which can preserve the semantic information by making sure the generated hash codes are also optimal for classification. Comparison experiments are conducted on two benchmark datasets to demonstrate the performance of our proposed approach.
Shuhan Qi, Xuan Wang 0002, Jian Guan 0001, Fengwei Jia, Lin Yao 0004
ICME3
2018 Position Paper on Recent Cybersecurity Trends: Legal Issues, AI and IoT
Yun-Ju Huang, Frankie Li, Jing Li 0045, Xuan Wang 0002, Yang Xiang 0001
NSS5
2018 Scalable graph based non-negative multi-view embedding for image ranking
Shuhan Qi, Xuan Wang 0002, Xuemeng Song, Zoe Lin Jiang
Neurocomputing2
2018 Practical attribute-based encryption: Outsourcing decryption, attribute revocation and policy updating
Zechao Liu, Zoe Lin Jiang, Xuan Wang 0002, Siu-Ming Yiu
J. Netw. Comput. Appl.3
2018 Polynomial dictionary learning algorithms in sparse representations
Jian Guan 0001, Xuan Wang 0002, Pengming Feng, Jing Dong 0001, Jonathon A. Chambers, Zoe Lin Jiang, Wenwu Wang 0001
Signal Process.2
2018 Securely Outsourcing ID3 Decision Tree in Cloud Computing
abstract
With the wide application of Internet of Things (IoT), a huge number of data are collected from IoT networks and are required to be processed, such as data mining. Although it is popular to outsource storage and computation to cloud, it may invade privacy of participants’ information. Cryptography‐based privacy‐preserving data mining has been proposed to protect the privacy of participating parties’ data for this process. However, it is still an open problem to handle with multiparticipant’s ciphertext computation and analysis. And these algorithms rely on the semihonest security model which requires all parties to follow the protocol rules. In this paper, we address the challenge of outsourcing ID3 decision tree algorithm in the malicious model. Particularly, to securely store and compute private data, the two‐participant symmetric homomorphic encryption supporting addition and multiplication is proposed. To keep from malicious behaviors of cloud computing server, the secure garbled circuits are adopted to propose the privacy‐preserving weight average protocol. Security and performance are analyzed.
Ye Li 0023, Zoe Lin Jiang, Xuan Wang 0002, En Zhang, Xianmin Wang
Wirel. Commun. Mob. Comput.3
2017 Leveraging Target-Oriented Information for Stance Classification
Jiachen Du, Ruifeng Xu 0001, Lin Gui 0003, Xuan Wang 0002
CICLing (2)4
2017 Semantic Video Carving Using Perceptual Hashing and Optical Flow
Guikai Xi, Zoe Lin Jiang, Siu-Ming Yiu, Liyang Yu, Xuan Wang 0002, Qi Han 0002, Qiong Li 0001
IFIP Int. Conf. Digital Forensics7
2017 Matrix of Polynomials Model Based Polynomial Dictionary Learning Method for Acoustic Impulse Response Modeling
abstract
We study the problem of dictionary learning for signals that can be represented as polynomials or polynomial matrices, such as convolutive signals with time delays or acoustic impulse responses.Recently, we developed a method for polynomial dictionary learning based on the fact that a polynomial matrix can be expressed as a polynomial with matrix coefficients, where the coefficient of the polynomial at each time lag is a scalar matrix.However, a polynomial matrix can be also equally represented as a matrix with polynomial elements.In this paper, we develop an alternative method for learning a polynomial dictionary and a sparse representation method for polynomial signal reconstruction based on this model.The proposed methods can be used directly to operate on the polynomial matrix without having to access its coefficients matrices.We demonstrate the performance of the proposed method for acoustic impulse response modeling.
Jian Guan 0001, Xuan Wang 0002, Pengming Feng, Jing Dong 0001, Wenwu Wang 0001
INTERSPEECH2
2017 Outsourced Privacy-Preserving Random Decision Tree Algorithm Under Multiple Parties for Sensor-Cloud Integration
Ye Li 0023, Zoe Lin Jiang, Xuan Wang 0002, Siu-Ming Yiu
ISPEC3
2017 Offline/online attribute-based encryption with verifiable outsourced decryption
abstract
Summary In this big data era, service providers tend to put the data in a third‐party cloud system. Social networking websites are typical examples. To protect the security and privacy of the data, data should be stored in encrypted form. This brings forth new challenges: how to allow different users to access only the authorized part of the data without decryption of the data. Attribute‐based encryption (ABE) offers fine‐grained access control policy over encrypted data such that users can decrypt successfully only if their attributes satisfy the policy. However, one drawback of ABE is that the computational cost grows linearly with the complexity of ciphertext policy or the number of attributes. The situation becomes worse for mobile devices with limited computing resources. To solve this problem, we adopt the offline/online technique combining with the verifiable outsourced computation technique to propose a new ciphertext‐policy ABE scheme using bilinear groups in prime order, supporting the offline/online key generation and encryption, as well as the verifiable outsourced decryption. As a result, most computations of key generation and encryption can be executed offline, and the majority of computational workload in decryption can be outsourced to third parties. The scheme is selectively chosen‐plaintext attack‐secure in the standard model. We also provide the proof of verifiability on outsourced decryption. The simulation results show that our proposed scheme can effectively reduce the computational cost imposed on resource‐constrained devices. Copyright © 2016 John Wiley & Sons, Ltd.
Zechao Liu, Zoe Lin Jiang, Xuan Wang 0002, Xinyi Huang 0001, Siu-Ming Yiu, Kunihiko Sadakane
Concurr. Comput. Pract. Exp.3
2017 Improving sentiment analysis via sentence type classification using BiLSTM-CRF and CNN
abstract
Different types of sentences express sentiment in very different ways. Traditional sentence-level sentiment classification research focuses on one-technique-fits-all solution or only centers on one special type of sentences. In this paper, we propose a divide-and-conquer approach which first classifies sentences into different types, then performs sentiment analysis separately on sentences from each type. Specifically, we find that sentences tend to be more complex if they contain more sentiment targets. Thus, we propose to first apply a neural network based sequence model to classify opinionated sentences into three types according to the number of targets appeared in a sentence. Each group of sentences is then fed into a one-dimensional convolutional neural network separately for sentiment classification. Our approach has been evaluated on four sentiment classification datasets and compared with a wide range of baselines. Experimental results show that: (1) sentence type classification can improve the performance of sentence-level sentiment analysis; (2) the proposed approach achieves state-of-the-art results on several benchmarking datasets.
Tao Chen 0026, Ruifeng Xu 0001, Yulan He 0001, Xuan Wang 0002
Expert Syst. Appl.4
2017 Clustering based virtual machines placement in distributed cloud computing
Xuan Wang 0002, Hejiao Huang
Future Gener. Comput. Syst.2
2017 Multiple level visual semantic fusion method for image re-ranking
Shuhan Qi, Fanglin Wang, Xuan Wang 0002, Jia Wei 0003, Jian Guan 0001
Multim. Syst.3
2017 Robust visual tracking via discriminative appearance model based on sparse coding
Hainan Zhao, Xuan Wang 0002
Multim. Syst.2
2017 A Two-Phase Improved Correlation Method for Automatic Particle Selection in Cryo-EM
abstract
Particle selection from cryo-electron microscopy (Cryo-EM) images is very important for high-resolution reconstruction of macromolecular structure. The methods of particle selection can be roughly grouped into two classes, template-matching methods and feature-based methods. In general, template-matching methods usually generate better results than feature-based methods. However, the accuracy of template-matching methods is restricted by the noise and low contrast of Cryo-EM images. Moreover, the processing speed of template-matching methods, restricted by the random orientation of particles, further limits their practical applications. In this paper, combining the advantages of feature-based methods and template-matching methods, we present a two-phase improved correlation method for automatic, fast particle selection. In Phase I, we generate a preliminary particle set using rotation-invariant features of particles. In Phase II, we filter the preliminary particle set using a correlation method to reduce the interference of the high noise background and improve the precision of particle selection. We apply several optimization strategies, including a modified adaboost algorithm, Divide and Conquer technique, cascade strategy and graphics processing unit parallel technique, to improve feature recognition ability and reduce processing time. In addition, we developed two correlation score functions for different correlation situations. Experimental results on the benchmark of Cryo-EM images show that our method can improve the accuracy and processing speed of particle selection significantly.
Fa Zhang 0001, Xuan Wang 0002, Zhiyong Liu 0002
IEEE ACM Trans. Comput. Biol. Bioinform.4
2017 Matryoshka Peek: Toward Learning Fine-Grained, Robust, Discriminative Features for Product Search
abstract
In sharp contrast to the traditional category/subcategory level image retrieval, product image search aims to find the images containing the exact same product. This is a challenging problem because in addition to being robust under different imaging conditions such as varying viewpoints and illumination changes, the features should also be able to distinguish the specific product among many similar products. Consequently, it is important to utilize a large dataset, containing many product classes, to learn a strongly discriminative representation. Building such a dataset requires laborious manual annotation. Toward learning fine-grained, robust, discriminative features for product image search, we present a novel paradigm that can construct the required dataset without any human annotation. Unlike other fine-grained recognition works that rely on high-quality annotated datasets and are very narrowly focused on a specific object category, our method handles multiple object classes and requires minimum human effort. First, an ImageNet pretrained model is used to generate product clusters. As the original features from ImageNet are not discriminative, the clusters generated by this unsupervised procedure contain much noise. We alleviate noise by explicitly modeling noise distribution and automatically detecting errors during learning. The proposed paradigm is general, requires minimum human efforts, and is applicable to any deep learning task where fine-grained discriminative features are desired. Extensive experiments on the ALISC dataset have demonstrated that our approach is sound and effective, surpassing the baseline GoogleNet model by 15.09%.
Zawlin Kyaw, Shuhan Qi, Ke Gao 0012, Hanwang Zhang, Jun Xiao 0001, Xuan Wang 0002, Tat-Seng Chua
IEEE Trans. Multim.7
2016 Generic Construction of Publicly Verifiable Predicate Encryption
abstract
There is an increasing trend for data owners to store their data in a third-party cloud server and buy the service from the cloud server to provide information to other users. To ensure confidentiality, the data is usually encrypted. Therefore, an encrypted data searching scheme with privacy preserving is of paramount importance. Predicate encryption (PE) is one of the attractive solutions due to its attribute-hiding merit. However, as cloud is not always trusted, verifying the searched results is also crucial. Firstly, a generic construction of Publicly Verifiable Predicate Encryption (PVPE) scheme is proposed to provide verification for PE. We reduce the security of PVPE to the security of PE. However, from practical point of view, to decrease the communication overhead and computation overhead, an improved PVPE is proposed with the trade-off of a small probability of error.
Chuting Tan, Zoe Lin Jiang, Xuan Wang 0002, Siu-Ming Yiu, Jin Li 0002, Yabin Jin
AsiaCCS3
2016 Key based data analytics across data centers considering bi-level resource provision in cloud computing
Lingmin Zhang, Hejiao Huang, Zoe Lin Jiang, Xuan Wang 0002
Future Gener. Comput. Syst.5
2016 Adaptive semi-supervised dimensionality reduction with sparse representation using pairwise constraints
Jia Wei 0003, Jiabing Wang, Qianli Ma 0001, Xuan Wang 0002
Neurocomputing5
2016 A joint appearance model of SRC and MFH for multi-objects tracking
Hainan Zhao, Shuhan Qi, Xuan Wang 0002
Neurocomputing3
2016 Resource provision algorithms in cloud computing: A survey
Hejiao Huang, Xuan Wang 0002
J. Netw. Comput. Appl.3
2016 Quality biased multimedia data retrieval in microblogs
Shuhan Qi, Peiguang Jing, Xuan Wang 0002, Liqiang Nie
J. Vis. Commun. Image Represent.3
2015 Dynamic Resource Provision for Cloud Broker with Multiple Reserved Instance Terms
Hejiao Huang, Xuan Wang 0002, Ding-Zhu Du
ICA3PP (1)4
2015 Learning Task Specific Distributed Paragraph Representations Using a 2-Tier Convolutional Neural Network
Tao Chen 0026, Ruifeng Xu 0001, Yulan He 0001, Xuan Wang 0002
ICONIP (1)4
2015 Outsourcing Two-Party Privacy Preserving K-Means Clustering Protocol in Wireless Sensor Networks
abstract
Nowadays wireless sensor network (WSN) is widely used in human-centric applications and environmental monitoring. Different institutes deploy their own WSNs for data collection and processing. It becomes a challenging problem when institutes collaborate to do data mining while intend to keep data privacy on each side. Privacy preserving data mining (PPDM) is used to solve the above problem, which enables multiple parties owning confidential data to run a data mining algorithm on their combined data, without revealing any unnecessary information to each other. However, due to the huge amount of data collected and the complexity of data mining algorithms, it is preferable to outsource most of the computations to the cloud. In this paper, we consider a scenario in which two parties with weak computational power need jointly run a k-means clustering protocol, at the same time outsource most of the computation of the protocol to the cloud. As a result, each party can have the correct result calculated by the data from both parties with most of the computation outsourced to the cloud. As for privacy, the data owned by one party should be kept confidential from both the other party and the cloud.
Zoe Lin Jiang, Siu-Ming Yiu, Xuan Wang 0002, Chuting Tan, Ye Li 0023, Zechao Liu, Yabin Jin
MSN4
2015 Live multimedia brand-related data identification in microblog
Shuhan Qi, Fanglin Wang, Xuan Wang 0002, Jia Wei 0003, Hainan Zhao
Neurocomputing3
2015 A Novel Arc Segmentation Approach for Document Image Processing
abstract
In document image processing, arc segmentation plays an important role in vectorization and graphic recognition. Moreover, the unsatisfactory results of several recent arc segmentation contests indicate that conventional methods are inadequate. This paper proposes a new arc segmentation algorithm called SymCAve (an acronym for Symmetry axis, Circle fitting and Average distribution points). First, we locate several seed points and adopt three strategies to ensure that the seed points are proper; then we calculate the center and radius utilizing the seed points. Second, the coordinates of the center and radius are adjusted by employing symmetry axes. Third, the average distribution points method is used to verify whether the points on the circumference are all black pixels. It is a complete circle if all of the points are black pixels. Otherwise, it is a partial circle if some of the points are black pixels and are continuous. Based on this information, the start and end angles of the partial circle can be determined. Finally, these arcs are verified to ensure that the results are accurate. Images and the evaluation tool were obtained from the GREC Workshop's Arc Segmentation contests, to test the systematic performance of the SymCAve algorithm. The experiments demonstrate that the proposed method can provide promising results. However, the algorithm has some drawbacks: it cannot detect a line with width of one pixel, small angles, and any large radius arcs. It is suited for segmenting images with appropriate symmetry axes.
Xuan Wang 0002, Kai Han 0001, Zoe Lin Jiang
Int. J. Pattern Recognit. Artif. Intell.2
2014 Fully Secure Ciphertext-Policy Attribute Based Encryption with Security Mediator
Yuechen Chen, Zoe Lin Jiang, Siu-Ming Yiu, Joseph K. Liu, Man Ho Au, Xuan Wang 0002
ICICS6
2014 SLA aware cost efficient virtual machines placement in cloud computing
abstract
Servers and network contribute about 60% to the total cost of data center in cloud computing. How to efficiently place virtual machines so that the cost can be saved as much as possible, while guaranteeing the quality of service plays a critical role in enhancing the competitiveness of service cloud provider. Considering the heterogeneous servers and the random property of multiple resources requirements of virtual machines, the problem is formulated as a multi-objective nonlinear programming in this paper. Virtual machine cluster with higher traffic is made staying together. This reduces the communication delay while saving the inter-server bandwidth consumption, especially the relatively scarce higher level bandwidth, by exploiting the topology information of data center. At the same time, statistic multiplex and newly defined “similarity” techniques are leveraged to consolidate virtual machines. The violation of resource capacity is kept at any designated minimal probability. Thus the quality of service will not be deteriorated while saving servers and network cost. An offline and an online algorithms are proposed to address this problem. Experiments compared with several baseline algorithms show the validity of the new algorithms: more cost is cut down at less computation effort.
Zhixiang He, Hejiao Huang, Xuan Wang 0002, Chonglin Gu, Lingmin Zhang
IPCCC4
2014 An Improved Correlation Method Based on Rotation Invariant Feature for Automatic Particle Selection
Xuan Wang 0002, Fa Zhang 0001
ISBRA4
2014 A Parallel Scheme for Three-Dimensional Reconstruction in Large-Field Electron Tomography
Jingrong Zhang, Fa Zhang 0001, Xuan Wang 0002, Zhiyong Liu 0002
ISBRA5
2014 Integrating local and global topological structures for semi-supervised dimensionality reduction
Jia Wei 0003, Qun-fang Zeng, Xuan Wang 0002, Jiabing Wang, Guihua Wen
Soft Comput.3
2013 A novel remote eye gaze tracking approach with dynamic calibration
abstract
Point of gaze estimation is the most important part of remote eye gaze tracking techniques. Though there are various point of gaze detection and estimation methods, few of them can satisfy the expectation of widely use. One primary reason is lack of accurate calibration method. In order to deal with this problem, an adaptive calibration technique is proposed based on cross-ratio. It can compensate estimation bias much better. Also, efficient remote eye gaze tracking parameters estimation procedures are employed in this paper. The eye gaze tracking approach in this paper allows users to have free head motion and has high accuracy rate, and it is based on infrared radiation (IR) light sources reflection on cornea. Experimental results demonstrate that considerable improvement can be achieved.
Kai Han 0001, Xuan Wang 0002, Hainan Zhao
MMSP2
2011 Improve Coreference Resolution with Parameter Tunable Anaphoricity Identification and Global Optimization
Shuhan Qi, Xuan Wang 0002
ICIC (3)2
2011 Using Hybrid Kernel Method for Question Classification in CQA
Shixi Fan, Xiaolong Wang 0001, Xuan Wang 0002, Xiaohong Yang
ICONIP (3)3
2011 Complex Detection Based on Integrated Properties
Lei Lin 0001, Chengjie Sun, Xiaolong Wang 0001, Xuan Wang 0002
ICONIP (1)5
2011 Diversifying Question Recommendations in Community-Based Question Answering
Yaoyun Zhang, Xiaolong Wang 0001, Xuan Wang 0002, Ruifeng Xu 0001, Buzhou Tang
ICONIP (3)3
2011 Diversifying Information Needs in Results of Question Retrieval
Yaoyun Zhang, Xiaolong Wang 0001, Xuan Wang 0002, Ruifeng Xu 0001, Jun Xu 0007, Shixi Fan
IJCNLP3
2010 Reranking for Stacking Ensemble Learning
Buzhou Tang, Qingcai Chen, Xuan Wang 0002, Xiaolong Wang 0001
ICONIP (1)3
2009 Protein Long Disordered Region Prediction Based on Profile-Level Disorder Propensities and Position-Specific Scoring Matrixes
abstract
Identification of long disordered regions in protein sequence is important for understanding protein function. In this work, a class of novel propensities at profile level is presented, namely, the order profile disorder propensities, which use the evolutionary information of profile for protein long disorder prediction. These propensities, combined with position-specific scoring matrices, are inputted to the logistic regression (LR) for the prediction of protein long disordered regions. In 5-fold cross-validation test, our method can achieve an area of 97.5% under the ROC cure. Testing on a blind-test set, our method is significantly more accurate than several existing disorder predictors.
Bin Liu 0014, Lei Lin 0001, Xiaolong Wang 0001, Xuan Wang 0002
BIBM4
2009 Improved Classification Based on Predictive Association Rules
abstract
Classification based on predictive association rules (CPAR) is a kind of association classification methods which combines the advantages of both associative classification and traditional rule-based classification. For rule generation, CPAR is more efficient than traditional rule-based classification because much repeated calculation is avoided and multiple literals can be selected to generate multiple rules simultaneously. Despite these advantages above in rule generation, the prediction processes have the weaknesses of class rule distribution imbalance and interruption of incorrect class rules. Further, it is useless to instances satisfying no rules. To tackle these problems, this paper presents Class Weighting Adjustment, Center Vector-based Pre-classification and Post-processing with Support Vector Machine. Experiments on Chinese text classification corpus TanCorp show that our algorithm achieves an average improvement of 5.91% on F1 score compared with CPAR.
Zhixin Hao, Xuan Wang 0002, Lin Yao 0004, Yaoyun Zhang
SMC2
2009 The Improvement of Q-learning Applied to Imperfect Information Game
abstract
There exist problems of slow convergence and local optimum in standard Q-learning algorithm. Truncated TD estimate returns efficiency and simulated annealing algorithm increase the chance of exploration. To accelerate the algorithm convergence speed and to avoid results in local optimum, this paper combines Q-learning algorithm, truncated TD estimation and simulated annealing algorithm. We apply improved Q-learning algorithm using into the imperfect information game (SiGuo military chess game), and realize a self-learning of imperfect information game system. Experimental outcomes show that this system can dynamically adjust each weight which describes game state according to the results. Further, it speeds up the process of learning, effectively simulates human intelligence and makes reasonable step, and significantly improves system performance.
Xuan Wang 0002, Lijiao Han, Jiajia Zhang 0001
SMC2
2009 CRF-based Active Learning for Chinese Named Entity Recognition
abstract
Conditional Random Fields (CRFs) have been used for many sequence labeling tasks and got excellent results. Further, the supervised model strongly depends on the huge training data. Active learning is a different way rather than relying on a large amount random sampling. However, random sampling constructively participates in the optimal choosing training examples. Based on different query strategies, active learning can combine with other machine learning methods to reduce the annotation cost while maintaining the accuracy. This paper proposes a new active learning strategy based on Information Density (ID) integrated with CRFs for Chinese Named Entity Recognition (NER). On Sighan bakeoff 2006 MSRA NER corpus, an F1 score of 77.2% is achieved by using only 10,000 labeled training sentences chosen by the proposed active learning strategy.
Lin Yao 0004, Chengjie Sun, Xiaolong Wang 0001, Xuan Wang 0002
SMC5
2009 Using Question Classification to Model User Intentions of Different Levels
abstract
User information need detection is a fundamental issue in automatic question answering systems. Based on real questions collected from on-line question answering communities, this paper proposes a three-level question type taxonomy to model user information need. The three levels are based on interrogative patterns, hidden user intentions and specific answer expectations. One question can have multiple types in level 2&3. Question type assignment of level 2&3 is subjective-orientated, and may vary between different users. Shallow lexical, syntactic and semantic features are used to model the inherent subjectivity of user intentions. Classification experiments are conducted on a corpus of real questions collected from the web. Different machine learning methods are employed. Experimental results are promising. This indicates the capability of modeling user information need and subjectivity statistically, and that strong correlations exist between question types of the same level.
Yaoyun Zhang, Xuan Wang 0002, Xiaolong Wang 0001, Shixi Fan, Daoxu Zhang
SMC2
2009 Prediction of protein binding sites in protein structures using hidden Markov support vector machine
abstract
BACKGROUND: Predicting the binding sites between two interacting proteins provides important clues to the function of a protein. Recent research on protein binding site prediction has been mainly based on widely known machine learning techniques, such as artificial neural networks, support vector machines, conditional random field, etc. However, the prediction performance is still too low to be used in practice. It is necessary to explore new algorithms, theories and features to further improve the performance. RESULTS: In this study, we introduce a novel machine learning model hidden Markov support vector machine for protein binding site prediction. The model treats the protein binding site prediction as a sequential labelling task based on the maximum margin criterion. Common features derived from protein sequences and structures, including protein sequence profile and residue accessible surface area, are used to train hidden Markov support vector machine. When tested on six data sets, the method based on hidden Markov support vector machine shows better performance than some state-of-the-art methods, including artificial neural networks, support vector machines and conditional random field. Furthermore, its running time is several orders of magnitude shorter than that of the compared methods. CONCLUSION: The improved prediction performance and computational efficiency of the method based on hidden Markov support vector machine can be attributed to the following three factors. Firstly, the relation between labels of neighbouring residues is useful for protein binding site prediction. Secondly, the kernel trick is very advantageous to this field. Thirdly, the complexity of the training step for hidden Markov support vector machine is linear with the number of training samples by using the cutting-plane algorithm.
Bin Liu 0014, Xiaolong Wang 0001, Lei Lin 0001, Buzhou Tang, Qiwen Dong, Xuan Wang 0002
BMC Bioinform.6
2008 Discriminative Learning of Syntactic and Semantic Dependencies
Shixi Fan, Xuan Wang 0002, Xiaolong Wang 0001
CoNLL3
2008 Chunking with Max-Margin Markov Networks
Buzhou Tang, Xuan Wang 0002, Xiaolong Wang 0001
PACLIC2
2008 Semantic Chunk Annotation for questions using Maximum Entropy
abstract
we present a ME (Maximum Entropy) model for Semantic Chunk Annotation in a Chinese Question and Answer (Q&A) system. The model was derived from a corpus of real world questions, which are collected from some discussion groups on the Internet. The questions are supposed to be answered by other people, so the questions are very complex. The semantic chunks were introduced. Feature for the model was described and MI (Mutual Information) was adopted for feature selection. The training data consists of 14000 sentences and the test data consists of 4000 sentences. The result: F-score is 90.68%.
Shixi Fan, Yaoyun Zhang, Wing W. Y. Ng, Xuan Wang 0002, Xiaolong Wang 0001
SMC4
2008 Avatars based Chinese Sign Language Synthesis System
abstract
A Chinese sign language synthesis system is proposed in this paper. H-Anim standard of VRML is adopted to describe the avatar. The traditional reverse kinematics resolution algorithm is employed and constraints of the elbow position are considered to make the editing of hand shapes and arm gestures efficient. Various interpolation methods are used for the motion of joints that have different degrees-of-freedom. The Chinese sign language synthesis system is implemented as a Web service which can be accessed by 3rd-party applications. It has been applied in sign language teaching and sign language translation.
Xuan Wang 0002, Lin Yao 0004, Deyuan Zhang, Hainan Zhao
SMC2
2008 A discriminative method for protein remote homology detection and fold recognition combining Top-n-grams and latent semantic analysis
abstract
BACKGROUND: Protein remote homology detection and fold recognition are central problems in bioinformatics. Currently, discriminative methods based on support vector machine (SVM) are the most effective and accurate methods for solving these problems. A key step to improve the performance of the SVM-based methods is to find a suitable representation of protein sequences. RESULTS: In this paper, a novel building block of proteins called Top-n-grams is presented, which contains the evolutionary information extracted from the protein sequence frequency profiles. The protein sequence frequency profiles are calculated from the multiple sequence alignments outputted by PSI-BLAST and converted into Top-n-grams. The protein sequences are transformed into fixed-dimension feature vectors by the occurrence times of each Top-n-gram. The training vectors are evaluated by SVM to train classifiers which are then used to classify the test protein sequences. We demonstrate that the prediction performance of remote homology detection and fold recognition can be improved by combining Top-n-grams and latent semantic analysis (LSA), which is an efficient feature extraction technique from natural language processing. When tested on superfamily and fold benchmarks, the method combining Top-n-grams and LSA gives significantly better results compared to related methods. CONCLUSION: The method based on Top-n-grams significantly outperforms the methods based on many other building blocks including N-grams, patterns, motifs and binary profiles. Therefore, Top-n-gram is a good building block of the protein sequences and can be widely used in many tasks of the computational biology, such as the sequence alignment, the prediction of domain boundary, the designation of knowledge-based potentials and the prediction of protein binding sites.
Bin Liu 0014, Xiaolong Wang 0001, Lei Lin 0001, Qiwen Dong, Xuan Wang 0002
BMC Bioinform.5
2007 Model fusion of conditional random fields
abstract
This paper introduces two model fusion methods on a series of sub-models of Conditional Random Fields (CRFs): majority voting and feature fusion. The former performs on the results of each participant without any consideration about the underlying details of each sub-model, and the latter takes place on feature level to produce modified feature weights of CRFs to merge all sub-models into a single one. Experiments on syntactic data and part-of-speech tagging problem shows that by dividing training corpus into small parts and using model fusion techniques, comparable results will be achieved.
Xuan Wang 0002, Yanbing Yu, Xiaolong Wang 0001
SMC2
2007 Intelligent chinese text input technology for mobile computing
abstract
A sentence-based Chinese text input method system is proposed in this paper, which is implemented on both Symbian S60 and Windows Mobile platform with such characters as easy-to-use, efficient and smart. The whole system is compacted within 150k, and can be integrated with cell phone, PDA and remoter.
Xuan Wang 0002, Lin Yao 0004, Xiaolong Wang 0001
SMC1
2006 A Maximum Entropy Approach to Chinese Pin Yin-To-Character Conversion
abstract
This paper introduces a new approach based upon maximum entropy (ME) frame to solve the Pinyin-to-character (PTC) conversation problem. Mostly there is more than one Chinese characters share the same Pinyin. The task of PTC algorithm is to distinguish such kind ambiguity. PTC can be regards as to classify a Pinyin to a special character according the context which is represented as feature in ME. By taking the advantage of ME, the local and non-local information are included, so the conversation performance is improved. Experiments show that 87% hit rate (without tone) is achieved.
Xuan Wang 0002, Lin Yao 0004, Waqas Anwar
SMC1
2005 A Hybrid Language Model Based On Statistics And Linguistic Rules
abstract
Language modeling is a current research topic in many domains including speech recognition, optical character recognition, handwriting recognition, machine translation and spelling correction. There are two main types of language models, the mathematical and the linguistic. The most widely used mathematical language model is the n-gram model inferred from statistics. This model has three problems: long distance restriction, recursive nature and partial language understanding. Language models based on linguistics present many difficulties when applied to large scale real texts. We present here a new hybrid language model that combines the advantages of the n-gram statistical language model with those of a linguistic language model which makes use of grammatical or semantic rules. Using suitable rules, this hybrid model can solve problems such as long distance restriction, recursive nature and partial language understanding. The new language model has been effective in experiments and has been incorporated in Chinese sentence input products for Windows and Macintosh OS.
Xiaolong Wang 0001, Daniel S. Yeung, James Nga-Kwok Liu, Robert Wing Pong Luk, Xuan Wang 0002
Int. J. Pattern Recognit. Artif. Intell.5