EDBT 2026 Demo / reviewers in the wild / expert
Jia Wang 0009
dblp:58/6299-9
· DBLP profile ↗
38ranked-venue papers
8as first author
25since 2021 · last 2026
0000-0002-3165-7051ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 2 first-author · 10 since 2021Databases, data management, data science and information retrieval · 10 · 1 first-author · 7 since 2021Systems, architecture and hardware · 8 · 2 first-author · 6 since 2021Computer networks · 4 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Security and privacy · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | From IDs to Semantics: A Generative Framework for Cross-Domain Recommendation with Adaptive Semantic TokenizationabstractCross-domain recommendation (CDR) is crucial for improving recommendation accuracy and generalization, yet traditional methods are often hindered by the reliance on shared user/item IDs, which are unavailable in most real-world scenarios. Consequently, many efforts have focused on learning disentangled representations through multi-domain joint training to bridge the domain gaps. Recent Large Language Model (LLM)-based approaches show promise, they still face critical challenges, including: (1) the \textbf{item ID tokenization dilemma}, which leads to vocabulary explosion and fails to capture high-order collaborative knowledge; and (2) \textbf{insufficient domain-specific modeling} for the complex evolution of user interests and item semantics. To address these limitations, we propose \textbf{GenCDR}, a novel \textbf{Gen}erative \textbf{C}ross-\textbf{D}omain \textbf{R}ecommendation framework. GenCDR first employs a \textbf{Domain-adaptive Tokenization} module, which generates disentangled semantic IDs for items by dynamically routing between a universal encoder and domain-specific adapters. Symmetrically, a \textbf{Cross-domain Autoregressive Recommendation} module models user preferences by fusing universal and domain-specific interests. Finally, a \textbf{Domain-aware Prefix-tree} enables efficient and accurate generation. Extensive experiments on multiple real-world datasets demonstrate that GenCDR significantly outperforms state-of-the-art baselines. Our code is available in the supplementary materials. Peiyu Hu, Wayne Lu, Jia Wang 0009 |
AAAI | 3 |
| 2026 | Breaking Down Market Barriers: Distilled Prompt-Tuning Approach for Cross-Market RecommendationabstractCross-market recommendation (CMR) faces severe challenges from distribution shifts between data-rich source markets and sparse target markets. Existing methods rely on a pre-training and fine-tuning paradigm for knowledge transfer, yet suffer from two key limitations: i) the objective gap between pre-training and full-parameter fine-tuning causes loss of generalized knowledge from source markets; ii) the high computational costs of extensive fine-tuning hinder scalability. To this end, we propose DCMPT, a novel Distilled Cross-Market Prompt-Tuning approach. DCMPT reframes the problem under a more efficient pre-training and prompt-tuning paradigm. Instead of full fine-tuning, we adapt a pre-trained universal backbone by freezing its weights and injecting a minimal set of learnable prompts to form a "student" model. To effectively optimize these prompts on sparse data, we introduce a novel teacher-student architecture: a specialized "teacher" model, trained exclusively on the target market, provides dense, market-specific supervision. This guidance is delivered via a dual distillation strategy designed to transfer global ranking patterns and adapt to local consumer tastes. Extensive experiments on real-world market datasets demonstrate that DCMPT significantly outperforms state-of-the-art methods, achieving superior target market performance with substantial parameter-efficiency. Leqi Zhang, Wayne Lu, Haiyang Zhang 0004, Elliott Wen, Zhixuan Liang, Jia Wang 0009 |
AAAI | 6 |
| 2026 | SRNeRV: A Scale-wise Recursive Framework for Neural Video Representation
Jia Wang 0009 |
ISCAS | 1 |
| 2026 | GNN-Based Item Indexing for LLM-Enhanced RecommendationabstractLarge language models (LLMs) have transformed recommender systems through strong semantic understanding and generalization. However, the design of item identifiers remains a critical bottleneck that directly affects recommendation quality. Traditional metadata-based identifiers introduce length variability and semantic ambiguity, whereas existing collaborative indexing (CID) approaches often neglect item attributes, show limited cross-dataset generalizability, and incur high computational cost at scale. To address these limitations, we propose a Graph Neural Network (GNN)–based item indexing framework with three coordinated innovations. First, we construct attribute-enriched co-occurrence graphs and use a GNN encoder to fuse item features with collaborative signals, yielding semantically informed representations that work well for attribute-rich catalogs. Second, we replace recursive spectral clustering with hierarchical agglomerative clustering on GNN embeddings, enabling direct control of index length via tree depth and reducing hyperparameter tuning across datasets. Third, we exploit localized message passing rather than global eigendecomposition, which provides considerably better runtime efficiency and is amenable to mini-batch training, supporting online index updates as interactions evolve. Across five benchmarks, GID achieves strong average ranking performance, showing larger improvements on sparse and attribute-rich datasets while remaining competitive in dense settings. The framework is robust under both seen and unseen prompt templates, which supports practical LLM-based recommendation. On sequential recommendation, GID improves HR@10 by 7.9% on average over the strongest baseline in each dataset. Senlin Mao, Ji Zhang 0001, Peng Zhang 0001, Ze Wang 0016, Xiaoyao Zheng, Jia Wang 0009 |
SIGIR | 6 |
| 2026 | Beyond shortcuts: Mitigating spurious correlations in radiological diagnosis with causal intervention
Xinyi Zeng, Jia Wang 0009, Yi Dong 0002, Wei Wang 0042, Yanji Jiang, Haiyang Zhang 0004 |
Knowl. Based Syst. | 2 |
| 2026 | Hierarchical Molecular Attention Network: Improving Molecular Property Prediction Through Substructure IdentificationabstractFew-shot molecular property prediction is a persisting challenge in many biology-related tasks, because the same molecule may exhibit different properties (e.g., active or inactive) in different tasks. Existing methods view all atoms as equally important and attend to predict the properties by averaging the features of similar molecules, which ignores key substructures within molecules and leads to poor prediction performance. Since a molecule implicitly includes key substructures, which determines the properties of the molecular, and the atom combination forms key substructures, we focus on the atom combination in this paper. With this, we propose the Hierarchical Molecular Attention Network (HMAN) to predict the molecular properties through combining atoms. First, we utilize the average pooling to extract both the prototype and the molecular features, and use Graph Neural Networks (GNN) to extract the atomic features, then concatenates the prototype and molecular features with all atomic features as input to the self-attention mechanism to calculate the different weights for atoms. Here, we select the top $B$ atoms with the highest attention scores to form the key substructures. Second, we view the formed key substructures as the query vectors and regard the molecular features as the key-value pairs, and then feed them into another self-attention to obtain the scores of key substructures. We still select $k$ key substructures with the highest scores, and weight the sum of a molecular feature and $k$ key substructures to predict the molecular properties. To train the HMAN, we design a new loss function, which includes Binary Cross-Entropy and a weighted negative log-likelihood. The former is to predict the molecular properties, and the last one is to optimize the weight distribution, rendering that the weight distribution matches the predicted probability distribution, with back-propagation. Theoretical analysis proves the convergence of the new loss function, and extensive experimental results demonstrate that HMAN significantly outperforms SOTA baseline models in molecular property prediction tasks. Liangzhe Chen, Xiaohui Cui, Haojun Zhu, Jia Wang 0009, Yizhang Jiang |
IEEE Trans. Comput. Biol. Bioinform. | 7 |
| 2025 | Customized Retrieval-Augmented Generation with LLM for Debiasing Recommendation UnlearningabstractModern recommender systems face a critical challenge in complying with privacy regulations like the “right to be forgotten“: removing a user's data without disrupting recommendations for others. Traditional unlearning methods address this by partial model updates, but introduce propagation bias-where unlearning one user's data distorts recommendations for behaviorally similar users, degrading system accuracy. While retraining eliminates bias, it is computationally prohibitive for large-scale systems. To address this challenge, we propose CRAGRU, a novel framework leveraging Retrieval-Augmented Generation (RAG) for efficient, user-specific unlearning that mitigates bias while preserving recommendation quality. CRAGRU decouples unlearning into distinct retrieval and generation stages. In retrieval, we employ three tailored strategies designed to precisely isolate the target user's data influence, minimizing collateral impact on unrelated users and enhancing unlearning efficiency. Subsequently, the generation stage utilizes an LLM, augmented with user profiles integrated into prompts, to reconstruct accurate and personalized recommendations without needing to retrain the entire base model. Experiments on three public datasets demonstrate that CRAGRU effectively unlearns targeted user data, significantly mitigating unlearning bias by preventing adverse impacts on non-target users, while maintaining recommendation performance comparable to fully trained original models. Our work highlights the promise of RAG-based architectures for building robust and privacy-preserving recommender systems. The source code is available at: https://github.com/zhanghaichao520/LLM_rec_unlearning. Chong Zhang 0006, Peiyu Hu, Jia Wang 0009 |
ICDM | 5 |
| 2025 | MMET: A Multi-Input and Multi-Scale Transformer for Efficient PDEs SolvingabstractPartial Differential Equations (PDEs) are fundamental for modeling physical systems, yet solving them in a generic and efficient manner using machine learning-based approaches remains challenging due to limited multi-input and multi-scale generalization capabilities, as well as high computational costs. This paper proposes the Multi-input and Multi-scale Efficient Transformer (MMET), a novel framework designed to address the above challenges. MMET decouples mesh and query points as two sequences and feeds them into the encoder and decoder, respectively, and uses a Gated Condition Embedding (GCE) layer to embed input variables or functions with varying dimensions, enabling effective solutions for multi-scale and multi-input problems. Additionally, a Hilbert curve-based reserialization and patch embedding mechanism decrease the input length. This significantly reduces the computational cost when dealing with large-scale geometric models. These innovations enable efficient representations and support multi-scale resolution queries for large-scale and multi-input PDE problems. Experimental evaluations on diverse benchmarks spanning different physical fields demonstrate that MMET outperforms SOTA methods in both accuracy and computational efficiency. This work highlights the potential of MMET as a robust and scalable solution for real-time PDE solving in engineering and physics-based applications, paving the way for future explorations into pre-trained large-scale models in specific domains. This work is open-sourced at https://github.com/YichenLuo-0/MMET. Jia Wang 0009, Dapeng Lan, Yu Liu 0011, Zhibo Pang |
IJCAI | 2 |
| 2025 | TCN-LSTM for Stock Prediction with Sentiment SignalsabstractThe stock market is a complex, nonlinear, and sentiment-driven system where price movements are influenced not only by quantitative indicators but also by market sentiment. While existing approaches—ranging from traditional time series models to deep learning—have made progress in modeling financial data, they often neglect the affective dimensions that drive investor behavior. To address this limitation, we propose a hybrid stock prediction model based on Temporal Convolutional Networks (TCN) and Long Short-Term Memory (LSTM), which captures both short-term fluctuations and long-term dependencies in time series data. Furthermore, we integrate sentiment features derived from financial news using natural language processing techniques, allowing the model to incorporate qualitative insights alongside numerical data. Experimental results across multiple stocks show that our model outperforms existing methods in terms of accuracy and robustness. This study demonstrates the value of combining affective computing with deep temporal modeling for more reliable and socially aware financial forecasting. Peiyu Hu, Ruoyu Liu, Jia Wang 0009 |
INDIN | 3 |
| 2025 | Role-aware Multi-agent Reinforcement Learning for Coordinated Emergency Traffic ControlabstractEmergency traffic control presents an increasingly critical challenge, requiring seamless coordination among emergency vehicles, regular vehicles, and traffic lights to ensure efficient passage for all vehicles. Existing models primarily only focus on traffic light control, leaving emergency and regular vehicles prone to delay due to the lack of navigation strategies. To address this issue, we propose the ***R*ole-aware *M*ulti-agent *T*raffic *C*ontrol (RMTC)** framework, which dynamically assigns appropriate roles to traffic components for better cooperation by considering their relations with emergency vehicles and adaptively adjusting their policies. Specifically, RMTC introduces a *Heterogeneous Temporal Traffic Graph (HTTG)* to model the spatial and temporal relationships among all traffic components (traffic lights, regular and emergency vehicles) at each time step. Furthermore, we develop a *Dynamic Role Learning* model to infer the evolving roles of traffic lights and regular vehicles based on HTTG. Finally, we present a *Role-aware Multi-agent Reinforcement Learning* approach that learns traffic policies conditioned on the dynamically roles. Extensive experiments across four public traffic scenarios show that RMTC outperforms existing traffic light control methods by significantly reducing emergency vehicle travel time, while effectively preserving traffic efficiency for regular vehicles. The code is released at [https://github.com/mingchenghexi/RMTC](https://github.com/mingchenghexi/RMTC). Hao Chen 0062, Jia Wang 0009, Senzhang Wang |
NeurIPS | 4 |
| 2025 | COFA: counterfactual attention framework for trustworthy wafer map failure classification
Kaiyue Feng, Jia Wang 0009, Chenke Yin, Andong Li |
Appl. Intell. | 2 |
| 2024 | Document Set Expansion with Positive-Unlabeled Learning Using Intractable Density EstimationabstractThe Document Set Expansion (DSE) task involves identifying relevant documents from large collections based on a limited set of example documents. Previous research has highlighted Positive and Unlabeled (PU) learning as a promising approach for this task. However, most PU methods rely on the unrealistic assumption of knowing the class prior for positive samples in the collection. To address this limitation, this paper introduces a novel PU learning framework that utilizes intractable density estimation models. Experiments conducted on PubMed and Covid datasets in a transductive setting showcase the effectiveness of the proposed method for DSE. Code is available from https://github.com/Beautifuldog01/Document-set-expansion-puDE. Haiyang Zhang 0004, Qiuyi Chen, Yanjie Zou, Jia Wang 0009, Yushan Pan, Mark Stevenson 0001 |
LREC/COLING | 4 |
| 2024 | Two-branch Network with Feature Fusion for Time Since Deposition Estimation of BloodstainsabstractIn bloodstain examination of collaborative medicine and forensics, the analysis and identification of time since deposition (TSD) plays a significant role. Traditional bloodstain analysis methods can only provide a rough estimate for the TSD of traces, and they are time-consuming. To address this issue, we propose a lightweight framework called Fourier Transform Infrared Network (FTIR-Net) that combines wavelet transform with deep learning. To be specific, we parallelly perform wavelet transform on infrared spectra and compute its second derivative to attain the sequential signal and spectral image. Then, the learning component employs two separate branches to extract features from the one-dimensional (1D) spectra signal and two-dimensional (2D) coefficient images provided by continuous wavelet transform (CWT). To effectively aggregate information from the spectral image, we design a Squeeze-and-Excitation Network (SENet) and combine it with 2D convolution. Finally, the extracted features are concatenated and flattened, followed by two fully connected (FC) layers for retention time analysis. Since the standard bloodstain dataset is lacking, we create a dataset that associates bloodstain with the attenuated total reflectance of Fourier transform infrared (ATR-FTIR). To demonstrate the effectiveness of our model in bloodstain analysis and exploit the properties of the proposed dataset, we present comprehensive experiments and ablation studies. Yushi Li, Yu Han 0001, Jia Wang 0009, Fangyu Wu 0001, Chenke Yin |
CSCWD | 4 |
| 2024 | How Pretrained Foundation Models and Cloud-Fog Automation Empower the Recycling of Electrical VehiclesabstractThe increasing prevalence of electric vehicles de-mands efficient and sustainable management of end-of-life lithium-ion batteries. This paper examines the use of Pretrained Foundation Models and Cloud-Fog Automation to improve robotic disassembly of these batteries. We evaluate the performance of two Vision Transformer Models, in tasks involving deformed, rusty, contaminated, and worn batteries. Our proposed architecture, utilizing cloud and fog computing, balances performance with resource efficiency, providing a scalable solution for electric vehicles battery recycling. Dapeng Lan, Jia Wang 0009, Dongxiao Hu, Zhibo Pang, Honghao Lyu |
INDIN | 3 |
| 2024 | Toward Multi-Agent Coordination in IoT via Prompt Pool-based Continual Reinforcement LearningabstractThe Internet of Things (IoT) represents a complex, dynamic environment where edge devices continuously optimize their policies to address a continual stream of tasks. Previous studies have typically relied on a rehearsal buffer containing data from past tasks or a known task identity to mitigate catastrophic forgetting. Our research, Prompt Pool-based Continual Reinforcement Learning (PPCRL), aims to create a more efficient memory system by expanding a single prompt into a prompt pool, allowing agents to automatically select a set of relevant prompts without needing task identity knowledge. Similar to prompt-based learning techniques, our approach utilizes a small trainable prompt pool to guide pre-trained models through sequential task learning systematically. This allows us to optimize prompts for guiding model predictions and effectively manage both shared and task-specific knowledge while maintaining model generalization. We conducted experiments on two multi-agent benchmarks where traditional methods suffer from significant performance degradation. In contrast, PPCRL demonstrates the capability to outperform baselines and exhibits high generalization ability. Chenhang Xu, Jia Wang 0009, Yong Yue 0001, Jun Qi 0001, Jieming Ma |
ISPA | 2 |
| 2024 | C²DR: Robust Cross-Domain Recommendation based on Causal DisentanglementabstractCross-domain recommendation aims to leverage heterogeneous information to transfers knowledge from a data-sufficient domain (source domain) to a data-scarce domain (target domain). Existing approaches mainly focus on learning single-domain user preferences and then employ a transferring module to obtain cross-domain user preferences, but ignore the modeling of users' domain specific preferences on items. We argue that incorporating domain-specific preferences from the source domain will introduce irrelevant information that fails to the target domain. Additionally, directly combining domain-shared and domain-specific information may hinder the target domain's performance. To this end, we propose C^2DR, a novel approach that disentangles domain-shared and domain-specific preferences from a causal perspective. Specifically, we formulate a causal graph to capture the critical causal relationships based on the underlying recommendation process, explicitly identifying domain-shared and domain-specific information as causal irrelevant variables. Then, we introduce disentanglement regularization terms to learn distinct representations of the causal variables that obey the independence constraints in the causal graph. Remarkably, our proposed method enables effective intervention and transfer of domain-shared information, thereby improving the robustness of the recommendation model. We evaluate the efficacy of C^2DR through extensive experiments on three real-world datasets, demonstrating significant improvements over state-of-the-art baselines. Menglin Kong, Jia Wang 0009, Yushan Pan, Haiyang Zhang 0004, Muzhou Hou |
WSDM | 2 |
| 2023 | The Dark Side of Explanations: Poisoning Recommender Systems with Counterfactual ExamplesabstractDeep learning-based recommender systems have become an integral part of several online platforms. However, their black-box nature emphasizes the need for explainable artificial intelligence (XAI) approaches to provide human-understandable reasons why a specific item gets recommended to a given user. One such method is counterfactual explanation (CF). While CFs can be highly beneficial for users and system designers, malicious actors may also exploit these explanations to undermine the system's security. Ziheng Chen 0002, Fabrizio Silvestri, Jia Wang 0009, Yongfeng Zhang 0005, Gabriele Tolomei |
SIGIR | 3 |
| 2023 | EID-GAN: Generative Adversarial Nets for Extremely Imbalanced Data AugmentationabstractImbalanced data cause deep neural networks to output biased results, and it becomes more serious when facing extremely imbalanced data regarding the outliers with tiny size (the ratio of the outlier size to the image size is around 0.05%). Many data argumentation models are proposed to supplement imbalanced data to alleviate biased results. However, the existing augmentation models cannot synthesize tiny outliers, which make the generated data unavailable. In this article, we propose a new augmentation model named extremely imbalanced data augmentation generative adversarial nets (EID-GANs) to address the extremely imbalanced data augmentation problem. First, we design a new penalty function by subtracting the outliers from the cropped region of generated instance to guide the generator to learn the features of outliers. After this, we combine the output value of the penalty function with the generator loss to jointly update the generator’s parameters with backpropagation. Second, we propose a new evaluation approach that adopts two outlier detectors withk-fold cross-validation to assess the availability of generated instances. We conduct extensive experiments to demonstrate the significant performance improvement of EID-GAN on two extremely imbalanced datasets, which are the industrial Piston and the Fabric datasets, and one general imbalanced dataset, i.e., the public DAGM dataset. The experimental results show that our EID-GAN outperforms the state-of-the-art (SOTA) augmentation models on different imbalanced datasets. Wei Li 0121, Jinlin Chen, Jiannong Cao 0001, Chao Ma 0008, Jia Wang 0009, Xiaohui Cui, Ping Chen 0001 |
IEEE Trans. Ind. Informatics | 5 |
| 2022 | ReLAX: Reinforcement Learning Agent Explainer for Arbitrary Predictive ModelsabstractCounterfactual examples (CFs) are one of the most popular methods for attaching post-hoc explanations to machine learning (ML) models. However, existing CF generation methods either exploit the internals of specific models or depend on each sample's neighborhood, thus they are hard to generalize for complex models and inefficient for large datasets. This work aims to overcome these limitations and introduces ReLAX, a model-agnostic algorithm to generate optimal counterfactual explanations. Specifically, we formulate the problem of crafting CFs as a sequential decision-making task and then find the optimal CFs via deep reinforcement learning (DRL) with discrete-continuous hybrid action space. Extensive experiments conducted on several tabular datasets have shown that ReLAX outperforms existing CF generation baselines, as it produces sparser counterfactuals, is more scalable to complex target models to explain, and generalizes to both classification and regression tasks. Finally, to demonstrate the usefulness of our method in a real-world use case, we leverage CFs generated by ReLAX to suggest actions that a country should take to reduce the risk of mortality due to COVID-19. Interestingly enough, the actions recommended by our method correspond to the strategies that many countries have actually implemented to counter the COVID-19 pandemic. Ziheng Chen 0002, Fabrizio Silvestri, Jia Wang 0009, He Zhu 0001, Hongshik Ahn, Gabriele Tolomei |
CIKM | 3 |
| 2022 | SecretHunter: A Large-scale Secret Scanner for Public Git RepositoriesabstractCollaborative software development platforms like GitHub have gained tremendous popularity. Unfortunately, many users have reportedly leaked authentication secrets (e.g., textual passwords and API keys) in public Git repositories and caused security incidents and finical loss. Recently, several tools were built to investigate the secret leakage in GitHub. However, these tools could only discover and scan a limited portion of files in GitHub due to platform API restrictions and band-width limitations. In this paper, we present SecretHunter, a real-time large-scale comprehensive secret scanner for GitHub. SecretHunter resolves the file discovery and retrieval difficulty via two major improvements to the Git cloning process. Firstly, our system will retrieve file metadata from repositories before cloning file contents. The early metadata access can help identify newly committed files and enable many bandwidth optimizations such as filename filtering and object deduplication. Secondly, SecretHunter adopts a reinforcement learning model to analyze file contents being downloaded and infer whether the file is sensitive. If not, the download process can be aborted to conserve bandwidth. We conduct a one-month empirical study to evaluate SecretHunter. Our results show that SecretHunter discovers 57% more leaked secrets than state-of-the-art tools. SecretHunter also reduces 85% bandwidth consumption in the object retrieval process and can be used in low-bandwidth settings (e.g., 4G connections). Elliott Wen, Jia Wang 0009, Jens Dietrich 0001 |
TrustCom | 2 |
| 2022 | GraphWare: A graph-based middleware enabling multi-robot cooperationabstractSummary Multi‐robot systems are widely used to handle complex and cooperative missions in various industrial applications. Although robotic middleware has become the key to reducing the complexity of multi‐robot application development, existing works still have limitations in controlling multiple robots to perform missions cooperatively. To enable multi‐robot cooperation, middleware should provide high‐level abstraction support, dynamic configuration, communication, and synchronization. In this article, we proposeGraphWare, a novel middleware that provides a graph‐based programming abstraction and its underlying runtime kernel for programming and building multi‐robot cooperation applications. The graph‐based programming abstraction can express cooperative missions without exposing the complexity of managing multiple robots. The runtime kernel configures and manages multiple heterogeneous robots to intelligently perform cooperative missions. We implementGraphWareand evaluate its performance with ball collection missions which are cooperatively accomplished by a group of mobile robots, and study the fault‐tolerance, flexibility, and scalability of the middleware in the realistic simulation. The experimental results demonstrate thatGraphWarefacilitates the multi‐robot cooperative mission with efficient mission completion time, high success rate, and marginal runtime overhead. Jinlin Chen, Jiannong Cao 0001, Zhixuan Liang, Zhiqin Cheng, Jia Wang 0009 |
Concurr. Comput. Pract. Exp. | 5 |
| 2022 | IRDA: Incremental Reinforcement Learning for Dynamic Resource AllocationabstractResource allocation problems often manifest as online decision-making tasks where the proper allocation strategy depends on the understanding of the allocation environment and resources workload. Most existing resource allocation methods are based on meticulously designed heuristics which ignore the patterns of incoming tasks, so the dynamics of incoming tasks cannot be properly handled. To address this problem, we mine the task patterns from the large volume of historical allocation data and propose a reinforcement learning model termed IRDA to learn the allocation strategy in an incremental way. We observe that historical allocation data is usually generated from the daily repeated operations, which is not independent and identically distributed. Training with partial of this dataset can make the allocation strategy converged already, thereby wasting a lot of remaining data. To improve the learning efficiency, we partition the whole historical allocation big dataset into multi-batch datasets, which forces the agent to continuously “explore” and learn on the distinct state spaces. IRDA reuses the strategy learned from the previous batch dataset and adapts it to the learning on the next batch dataset, so as to incrementally learn from multi-batch datasets and improve the allocation strategy. We apply the proposed method to handle baggage carousel allocation at Hong Kong International Airport (HKIA). The experimental results show that IRDA is capable of incrementally learning from multi-batch datasets, and improves the baggage carousel resource utilization by around 51.86 percent compared to the current baggage carousel allocation system at HKIA. Jia Wang 0009, Jiannong Cao 0001, Senzhang Wang, Zhongyu Yao, Wengen Li |
IEEE Trans. Big Data | 1 |
| 2021 | Spring Buddy: A Self-Adaptive Elastic Memory Management Scheme for Efficient Concurrent Allocation/Deallocation in Cloud Computing SystemsabstractWithin the cloud computing scenario, each server usually carries multiple service processes, which intensifies the concurrency pressure of the system. As a result, the process of memory management during page allocation and deallocation becomes a significant bottleneck. Although several methods such as Buddy System and Inverse Buddy System (iBuddy) have been proposed to improve the performance of memory management, they cannot adapt to the highly concurrent environment of cloud computing, because they either force the memory allocation/deallocation requests to be serialized or bring extra fragmentation. To address the above problem, we propose Spring Buddy, which improves the concurrency of both memory allocation and deallocation and avoids unnecessary fragmentation. It can detect the changes of system- and process-level memory request patterns and dynamically adjust the organization of page frames. Inventively, Spring Buddy uses the spring core layer to provide both concurrent response and resource aggregation capability which is adapted to the system's concurrency pressure, and also uses the spring lazy layer to further mitigate the system resource contention through process behavior prediction. To demonstrate the effectiveness of Spring Buddy, we implement it in the Linux kernel. The results demonstrate that Spring Buddy can reduce memory allocation latency by 71.47 % and deallocation latency by 93.20% on average compared to the existing methods. Yihui Lu, Chentao Wu, Jia Wang 0009, Xiaoming Gao, Jie Li 0002, Minyi Guo |
ICPADS | 4 |
| 2021 | CANE: community-aware network embedding via adversarial training
Jia Wang 0009, Jiannong Cao 0001, Wei Li 0121, Senzhang Wang |
Knowl. Inf. Syst. | 1 |
| 2021 | Learning Graph Representation With Generative Adversarial NetsabstractGraph representation learning aims to embed each vertex in a graph into a low-dimensional vector space. Existing graph representation learning methods can be classified into two categories: generative models that learn the underlying connectivity distribution in a graph, and discriminative models that predict the probability of edge between a pair of vertices. In this paper, we propose GraphGAN, an innovative graph representation learning framework unifying the above two classes of methods, in which the generative and the discriminative model play a game-theoretical minimax game. Specifically, for a given vertex, the generative model tries to fit its underlying true connectivity distribution over all other vertices and produces “fake” samples to fool the discriminative model, while the discriminative model tries to detect whether the sampled vertex is from ground truth or generated by the generative model. With the competition between these two models, both of them can alternately and iteratively boost their performance. Moreover, we propose a novel graph softmax as the implementation of the generative model to overcome the limitations of traditional softmax function, which can be proven satisfying desirable properties of normalization, graph structure awareness, and computational efficiency. Through extensive experiments on real-world datasets, we demonstrate that GraphGAN achieves substantial gains in a variety of applications, including graph reconstruction, link prediction, node classification, recommendation, and visualization, over state-of-the-art baselines. Hongwei Wang 0004, Jia Wang 0009, Miao Zhao, Weinan Zhang 0001, Wenjie Li 0002, Xing Xie 0001, Minyi Guo |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2020 | Recursive Balanced k-Subset Sum Partition for Rule-constrained Resource AllocationabstractBalanced rule-constrained resource allocation aims to evenly distribute tasks to different processors under allocation rule constraints. Conventional heuristic approach fails to achieve optimal solution while simple brute force method has the defect of high computational complexity. To address these limitations, we propose recursive balanced k-subset sum partition (RBkSP), in which iterative 'cut-one-out' policy is employed that in each round, only one subset whose weight of tasks sums up to 1/k of the total weight of all tasks is taken out from the set. In a single partition, we first create a dynamic programming table with its elements recursively computed, then use 'zig-zag search' method to explore the table, find out elements with optimal subset partition and assign different partitions to proper places. Next, to resolve conflicts during allocation, we use simple but effective heuristic method to adjust the allocation of tasks that is contradicted to allocation rules. Testing results show RBkSP can achieve more balanced results with lower computational complexity over classical benchmarks. Zhuo Li 0010, Jiannong Cao 0001, Zhongyu Yao, Wengen Li, Yu Yang 0012, Jia Wang 0009 |
CIKM | 6 |
| 2020 | BigARM: A Big-Data-Driven Airport Resource Management Engine and Application Tools
Ka-Ho Wong, Jiannong Cao 0001, Yu Yang 0012, Wengen Li, Jia Wang 0009, Zhongyu Yao, Suyan Xu, Esther Ahn Chian Ku, Chun On Wong, David Leung |
DASFAA (3) | 5 |
| 2020 | Joint Topic-Semantic-aware Social Matrix Factorization for online voting recommendation
Jia Wang 0009, Hongwei Wang 0004, Miao Zhao, Jiannong Cao 0001, Zhuo Li 0010, Minyi Guo |
Knowl. Based Syst. | 1 |
| 2019 | Pattern-RL: Multi-robot Cooperative Pattern Formation via Deep Reinforcement LearningabstractAutonomous, arbitrary pattern formation is one of the most critical applications in multi-robot systems, where robots are required to form into circles, lines, and meshes or any other desired configuration. This task is important in military applications, search and rescue operations, and visual inspection of infrastructure and equipment tasks to name a few. Most existing works are very rigid, and only able to form certain shapes, where slight target changes can cause failure in the predefined pattern-specific rules and trigger algorithm redesign. We propose a novel, deep reinforcement learning-based method that generates general-purpose pattern formation strategies, in the form of deep neural networks (DNN), for any target pattern. Our method uses the trial-and-error feedback of each round of training to gradually generate the pattern formation strategy. Thus, robots are able to query the trained DNN model to select their optimal directions and speeds in a fully distributed manner. Considering that reinforcement learning models do not perform well with large state spaces and highly variant training samples, we employ auto-encoders to learn the condensed representation for each state and compute model-free policy gradients for arbitrary pattern formation. We experimentally show that groups of robots are able to form various general target patterns while minimizing the number of completion time steps. Jia Wang 0009, Jiannong Cao 0001, Milos Stojmenovic, Miao Zhao, Jinlin Chen, Shan Jiang 0005 |
ICMLA | 1 |
| 2019 | Decentralized Algorithm for Repeating Pattern Formation by Multiple RobotsabstractRecently, much attention is paid to multi-robot systems due to their widespread applications such as warehouse robotics, persistent surveillance, and exploration of unknown environments. Although urgently required by the applications, coordination among multiple robots remains to be challenging. Among the problems of multi-robot coordination, pattern formation serves a fundamental one. It aims to control a group of robots to form a desired shape with some certain goals such as best formation quality, minimum makespan or minimum total distance. Existing works mainly focus on the formation of certain patterns, such as repeating squares or a circle. those approaches cannot be generalized to arbitrary pattern formation. In this paper, we propose a decentralized algorithm for a multi-robot system to generate a given formation with an arbitrary repeating pattern. We introduce basic pattern graph and assembling graph to define a repeating pattern and formation quality for measurement. Towards solving the repeating pattern formation problem, our approach is divided into two phases. The robots are grouped into multiple basic patterns in the first phase, and the patterns are assembled level by level in the second phase. Simulations and real-world experiments indicate the effectiveness and practicability of our approach. Shan Jiang 0005, Junbin Liang, Jiannong Cao 0001, Jia Wang 0009, Jinlin Chen, Zhixuan Liang |
ICPADS | 4 |
| 2019 | NFVdeep: adaptive online service function chain deployment with deep reinforcement learningabstractWith the evolution of network function virtualization (NFV), diverse network services can be flexibly offered as service function chains (SFCs) consisted of different virtual network functions (VNFs). However, network state and traffic typically exhibit unpredictable variations due to stochastically arriving requests with different quality of service (QoS) requirements. Thus, an adaptive online SFC deployment approach is needed to handle the real-time network variations and various service requests. In this paper, we firstly introduce a Markov decision process (MDP) model to capture the dynamic network state transitions. In order to jointly minimize the operation cost of NFV providers and maximize the total throughput of requests, we propose NFVdeep, an adaptive, online, deep reinforcement learning approach to automatically deploy SFCs for requests with different QoS requirements. Specifically, we use a serialization-and-backtracking method to effectively deal with large discrete action space. We also adopt a policy gradient based method to improve the training efficiency and convergence to optimality. Extensive experimental results demonstrate that NFVdeep converges fast in the training process and responds rapidly to arriving requests especially in large, frequently transferred network state space. Consequently, NFVdeep surpasses the state-of-the-art methods by 32.59% higher accepted throughput and 33.29% lower operation cost on average. Yikai Xiao, Qixia Zhang, Fangming Liu, Jia Wang 0009, Miao Zhao |
IWQoS | 4 |
| 2018 | GraphGAN: Graph Representation Learning With Generative Adversarial NetsabstractThe goal of graph representation learning is to embed each vertex in a graph into a low-dimensional vector space. Existing graph representation learning methods can be classified into two categories: generative models that learn the underlying connectivity distribution in the graph, and discriminative models that predict the probability of edge existence between a pair of vertices. In this paper, we propose GraphGAN, an innovative graph representation learning framework unifying above two classes of methods, in which the generative model and discriminative model play a game-theoretical minimax game. Specifically, for a given vertex, the generative model tries to fit its underlying true connectivity distribution over all other vertices and produces "fake" samples to fool the discriminative model, while the discriminative model tries to detect whether the sampled vertex is from ground truth or generated by the generative model. With the competition between these two models, both of them can alternately and iteratively boost their performance. Moreover, when considering the implementation of generative model, we propose a novel graph softmax to overcome the limitations of traditional softmax function, which can be proven satisfying desirable properties of normalization, graph structure awareness, and computational efficiency. Through extensive experiments on real-world datasets, we demonstrate that GraphGAN achieves substantial gains in a variety of applications, including link prediction, node classification, and recommendation, over state-of-the-art baselines. Hongwei Wang 0004, Jia Wang 0009, Miao Zhao, Weinan Zhang 0001, Xing Xie 0001, Minyi Guo |
AAAI | 2 |
| 2018 | Performance Optimization on Dynamic Adaptive Streaming over HTTP in Multi-User MIMO LTE NetworksabstractRecent years have witnessed the emergence and ongoing proliferation of dynamic adaptive streaming over HTTP (DASH1), which reuses web servers with HTTP communication instead of relying on RTSP/RTP/RTCP-based media server and promises to be capable of automatically tuning to bandwidth dynamics. Aware of its excellent performance, the third generation partnership project (3GPP) long term evolution (LTE) has adopted DASH (with specific codecs and operating modes) for use over mobile wireless networks in order to realize ubiquitous multimedia delivery. In a multi-user multiple-input-multiple-output (MU-MIMO) LTE system, spatial multiplexing gain can be achieved by making sure the transmitter to deliver distinct data streams to multiple receivers simultaneously, which provides the choices to opportunistically schedule the preferred receivers each time for a common time-frequency resource. In such a system, one of the major challenges to enhance DASH performance is to design an effective scheduler that can fully enjoy the benefit of spatial reuse as well as guaranteeing satisfactory video services for all users. To this end, in this paper, we propose a utility maximization framework (UMF) for DASH application delivered over MU-MIMO LTE downlinks. In particular, we characterize DASH performance by a combined utility function in terms of average video rate, playback buffer status, and battery energy state. Correspondingly, we develop a utility-based scheduler that selects multiple user equipments (UEs) to share each common network resource under the consideration of precoding-based MU-MIMO links in order to maximize system-wide DASH performance. We prove the NP-hardness of the scheduling problem and propose a priority search algorithm to provide time-efficient solution. We further incorporate novel rate adaptation on the application layer for the scheduled UEs to dynamically set the requested encoding bitrates to explore the balance between agile responsiveness and shifting smoothness. Extensive system-level simulations with realistic video trace validate the effectiveness of our framework in terms of rate adaptability, playback buffer depletion percentage, and battery energy saving. Miao Zhao, Jia Wang 0009, Mingquan Wu, Hong Heather Yu |
IEEE Trans. Mob. Comput. | 3 |
| 2017 | Joint Topic-Semantic-aware Social Recommendation for Online VotingabstractOnline voting is an emerging feature in social networks, in which users can express their attitudes toward various issues and show their unique interest. Online voting imposes new challenges on recommendation, because the propagation of votings heavily depends on the structure of social networks as well as the content of votings. In this paper, we investigate how to utilize these two factors in a comprehensive manner when doing voting recommendation. First, due to the fact that existing text mining methods such as topic model and semantic model cannot well process the content of votings that is typically short and ambiguous, we propose a novel Topic-Enhanced Word Embedding (TEWE) method to learn word and document representation by jointly considering their topics and semantics. Then we propose our Joint Topic-Semantic-aware social Matrix Factorization (JTS-MF) model for voting recommendation. JTS-MF model calculates similarity among users and votings by combining their TEWE representation and structural information of social networks, and preserves this topic-semantic-social similarity during matrix factorization. To evaluate the performance of TEWE representation and JTS-MF model, we conduct extensive experiments on real online voting dataset. The results prove the efficacy of our approach against several state-of-the-art baselines. Hongwei Wang 0004, Jia Wang 0009, Miao Zhao, Jiannong Cao 0001, Minyi Guo |
CIKM | 2 |
| 2017 | Uniform Circle Formation by Asynchronous Robots: A Fully-Distributed ApproachabstractRecent advances in robotics technology have made it practical to deploy a large number of inexpensive robots in a wide range of application domains. In many of those applications, a group of autonomous robots is required to form a predefined geometric shape such as a line or a circle. This problem, namely pattern formation problem, is one of the most important coordination problems in multi-robot systems. A particular pattern extensively studied in literature is the uniform circle, and the corresponding problem is called uniform circle formation. In uniform circle formation, a set of simple mobile robots (asynchronous, autonomous), starting from arbitrary positions on the plane, have to arrange themselves on the vertices of a regular polygon eventually. Towards addressing the problem, existing works usually make conveniently strong assumptions, i.e., the robots are regarded as mass points and have unlimited sensing and communication range. The question of whether the robots with actual size and limited sensing and communication range could form a uniform circle, to our knowledge, has remained open. In this paper, we propose a new approach towards addressing this issue. Three phases, consensus on the circle, circle formation, and uniform transformation, constitute our approach. Inside our approach, there are some new distributed algorithms such as convex hull construction and cardinality estimation. Simulation result, theoretical analysis, and successful deployment have shown the effectiveness and practicability of our approach. Shan Jiang 0005, Jiannong Cao 0001, Jia Wang 0009, Milos Stojmenovic, Julien Bourgeois |
ICCCN | 3 |
| 2017 | Fault-Tolerant Pattern Formation by Multiple Robots: A Learning ApproachabstractIn the field of multi-robot system, the problem of pattern formation has attracted considerable attention. However, the faulty sensor input of each robot is crucial for such system to act reliably in practice. Existing works focus on assuming certain noise model and reducing the noise impact. In this work, we propose to use a learning-based method to overcome this kind of barrier. By interacting with the environment, each robot learns to adapt its behavior to eliminate the malfunctions in the sensors and the actuators. Moreover, we plan to evaluate the proposed algorithms by deploying it into the multi-robot platform developed in our research lab. Jia Wang 0009, Jiannong Cao 0001, Shan Jiang 0005 |
SRDS | 1 |
| 2017 | Wi-friend: Identifying potential real life friends nearbyabstractNowadays, with the help of various on-line social networking applications, one can freely interact with any person in their friends list. However, in most conditions, people in your friends list are those you know a priori, either via face-to-face talking or introducing by a common friend. This way of establishing ones friends list can miss many potential friends nearby. These `potential real life friends nearby are those who are invisible to you (you do not know before), physically close to you (e.g. taking the same subway for commuting, doing exercises at the same time slot), and share the same interest with you. Correspondingly, we developed Wi-friend. Wi-friend, once installed in your mobile phone, can help to find these `potential friends nearby'. Using Wi-friend, one only needs to briefly specify his/her interest. Then the smartphone starts to look for nearby people with the same interest and add each other into ones friends list automatically. One special feature of Wi-friend is that it does not rely on Internet connection. We leverage the Wi-Fi tethering technique, and let ones smartphone to switch between Wi-Fi hot-spot mode and Wi-Fi client mode to exchange information with nearby people. In addition, Wi-friend can dynamically determine how long to stay in each mode to maximize the number of people who can exchange information. Extensive experiments and simulations demonstrated the technical feasibility and the effectiveness of our Wi-friend system. Jia Wang 0009, Xuefeng Liu 0001, Jiannong Cao 0001 |
WoWMoM | 1 |
| 2015 | Industry-friendly engineering tools for wireless home automation devicesabstractAlthough home automation (HA) systems in the wired domain are widely accepted by consumers, in today's industry, the mega trend is steering HA systems along a wireless way. Theoretically, wireless solutions are able to provide HA systems with more flexibility and thus reducing engineering costs. In practice, however, deploying wireless HA systems actually requires more costs and efforts due to the lack of versatile software tools to support the whole engineering process. This paper defines and evaluates the engineering workflow and architecture for home automation systems. The proposed architecture is studied and implemented based on web technologies and graphical configuration environments, with the aim of reducing workloads of HA engineers at every stage. A prototype has been implemented to demonstrate the technical feasibility of the proposed architecture. Jia Wang 0009, Zhibo Pang, Valeriy Vyatkin |
INDIN | 1 |