EDBT 2026 Demo / reviewers in the wild / expert
Jianxin Liao
dblp:60/4951
· DBLP profile ↗
241ranked-venue papers
22as first author
150since 2021 · last 2026
0000-0003-1486-0573ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 90 · 12 first-author · 45 since 2021Artificial intelligence and machine learning · 70 · 55 since 2021Graphics, computer vision, multimedia, augmented reality and games · 53 · 3 first-author · 40 since 2021Databases, data management, data science and information retrieval · 19 · 1 first-author · 11 since 2021Systems, architecture and hardware · 18 · 15 since 2021Software engineering, systems software and programming languages · 14 · 3 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 3 since 2021Security and privacy · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Bridging the Tokenizer Gap: Semantics and Distribution-aware Knowledge Transfer for Unbiased Cross-Tokenizer DistillationabstractCross-tokenizer knowledge distillation, where the teacher and student employ different tokenizers, is becoming increasingly prevalent, yet it poses underexplored challenges: existing methods fail to capture the rich knowledge encoded in teacher logits, as evidenced by the neglect of semantic information, inaccurate and biased logit alignment, and discarding distributional structure—ultimately leading to unfavorable distillation. To address these issues, we propose SeDi, a semantics and distribution-aware knowledge transfer framework tailored for cross-tokenizer distillation. To preserve factual knowledge, SeDi employs bipartite graph-based alignment at the tokenization level and a sliding window re-encoding strategy at the vocabulary level, enabling unbiased transfer of the teacher’s next-token predictions into the student’s vocabulary space. To further retain distributional information, we align the student’s entropy with that of the teacher by incorporating the student’s own logits during training, which helps to mitigate the exposure bias problem. Experiments on ten datasets across three task domains and five different teacher-student model pairs with varying vocabulary sizes demonstrate that SeDi delivers substantial improvements, with gains of up to 19.8%. Huazheng Wang, Yongcheng Jing, Haifeng Sun 0001, Jingyu Wang 0001, Jianxin Liao, Leszek Rutkowski, Dacheng Tao |
AAAI | 5 |
| 2026 | Example Quality Matters: Multi-Aspects Example Augmentation for Private Library ProgrammingabstractYuhao Li, Haifeng Sun, Xuesong Zhang, Shu Yao, Haoyu Zheng, Yvchuan Wang, Huazheng Wang, Zirui Zhuang, Qi Qi, Jianxin Liao, Jingyu Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Haifeng Sun 0001, Shu Yao, Yvchuan Wang, Huazheng Wang, Zirui Zhuang, Qi Qi 0001, Jianxin Liao, Jingyu Wang 0001 |
ACL (1) | 10 |
| 2026 | Erasing Without Remembering: Implicit Knowledge Forgetting in Large Language ModelsabstractHuazheng Wang, Yongcheng Jing, Haifeng Sun, Yingjie Wang, Jingyu Wang, Jianxin Liao, Dacheng Tao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Huazheng Wang, Yongcheng Jing, Haifeng Sun 0001, Jingyu Wang 0001, Jianxin Liao, Dacheng Tao |
ACL (1) | 6 |
| 2026 | From Generation to Guarantee: Intent-Based Configuration Update with Verification Feedback
Lingqi Guo, Qi Qi 0001, Haifeng Sun 0001, Yuxing Peng 0007, Zirui Zhuang, Bo He 0003, Shaoling Sun, Jianxin Liao, Jingyu Wang 0001 |
INFOCOM | 9 |
| 2026 | RecFlow: Unlocking GPU Efficiency for DLRM Inference via Fine-Grained Parallelism and Incremental Batching
Siheng Pan, Shaolong Li, Minwei Zhang, Shuxi Guo, Haifeng Sun 0001, Qi Qi 0001, Zirui Zhuang, Jianxin Liao, Jing Wang 0039 |
INFOCOM | 10 |
| 2026 | Ubi Grip: Ubiquitous Grip-Based Tangible Object Utilization in Augmented RealityabstractTangible Augmented Reality (AR) enhances user immersion in virtual world by providing haptic feedback through physical proxy objects. However, existing approaches primarily focus on selecting proxy objects based on their global physical properties, neglecting the utilization of local features. Besides, the prevailing strategy of mapping one virtual object to a single dedicated physical proxy creates an inherent switching cost, limiting flexibility and efficiency. Additionally, due to challenges such as real-time performance, generalization and occlusion, the vision-based hand-object tracking remains a difficult task. In this paper, we propose Ubi Grip, an universal hand-object interaction framework for creating grip-based tangible AR applications based on the local graspable feature and a comprehensive hand-object interaction attributes methodology. We employ a lightweight object tracking method to perform tracking, utilizing a hand mask filter and transformation strategy to optimize object pose based on the hand-held properties. Moreover, we design a user-defined workflow for grasping tangible objects, allowing users to switch grips and map interactions. We evaluated our system through comprehensive algorithmic benchmarks and a user study. The benchmarks demonstrate our SOTA performance in object pose estimation and generalization, while the user study validates system usability, providing deeper insights. Xufeng Jian, Guangtian Liu, Xiayang Zhou, Haifeng Sun 0001, Qi Qi 0001, Pengfei Ren 0001, Shan Jiang 0008, Jing Wang 0039, Jianxin Liao, Jingyu Wang 0001 |
VR | 11 |
| 2026 | MoCS: Modular configuration synthesis via large language models and graph neural network-augmented recommendation
Yuqi Dai, Hua Zhang 0002, Jingyu Wang 0001, Jianxin Liao |
Comput. Commun. | 4 |
| 2026 | RNV-RL: Relational Network Verification using Reinforcement Learning
Yuqi Dai, Hua Zhang 0002, Jingyu Wang 0001, Jianxin Liao |
J. Netw. Comput. Appl. | 4 |
| 2026 | OIPR: Evaluation for Time-Series Anomaly Detection Inspired by Operator InterestabstractWith the growing adoption of time-series anomaly detection (TAD) technology, numerous studies have employed deep learning-based detectors to analyze time-series data in the fields of Internet services, industrial systems, and sensors. The selection and optimization of anomaly detectors strongly rely on the availability of an effective evaluation for TAD performance. Since anomalies in time-series data often manifest as a sequence of points, conventional metrics that solely consider the detection of individual points are inadequate. Existing TAD evaluators typically employ point-based or event-based metrics to capture the temporal context. However, point-based evaluators tend to overestimate detectors that excel only in detecting long anomalies, while event-based evaluators are susceptible to being misled by fragmented detection results. To address these limitations, we propose OIPR1, a novel TAD evaluator with area-based metrics. It models the process of operators receiving detector alarms and handling anomalies, utilizing area under the operator interest curve to evaluate TAD performance. Furthermore, we build a special scenario dataset to compare the characteristics of different evaluators. Through experiments conducted on the special scenario dataset and five real-world datasets, we demon-strate the remarkable performance of OIPR in extreme and complex scenarios. It achieves a balance between point and event perspectives, overcoming their primary limitations and offering applicability to broader situations. Yuhan Jing, Jingyu Wang 0001, Lei Zhang 0094, Haifeng Sun 0001, Bo He 0003, Zirui Zhuang, Chengsen Wang, Qi Qi 0001, Jianxin Liao |
IEEE Trans. Dependable Secur. Comput. | 9 |
| 2026 | HyperWay: Proactively Mitigating Transient Congestion With Edge Capsule Tunnel in Massive IoT
Bo He 0003, Jinsheng Zhang, Jingyu Wang 0001, Qi Qi 0001, Haifeng Sun 0001, Zirui Zhuang, Yuhan Jing, Jing Shang 0001, Jianxin Liao |
IEEE Trans. Mob. Comput. | 10 |
| 2026 | LLM-Powered Intent-Driven Configuration Generation for Multi-Vendor Networks
Jingyu Wang 0001, Bo He 0003, Jinyu Zhao, Yixin Xuan, Haifeng Sun 0001, Qi Qi 0001, Junzhe Liang, Zirui Zhuang, Jianxin Liao |
IEEE Trans. Netw. Serv. Manag. | 9 |
| 2026 | Region Partitioning-Based Scalable Real-Time Network Verification via Native Distributed Architecture
Bo He 0003, Lingqi Guo, Chenyang Zhao 0005, Qi Qi 0001, Zirui Zhuang, Haifeng Sun 0001, Gong Zhang 0001, Jianxin Liao, Cheng Huang 0001, Jingyu Wang 0001 |
IEEE Trans. Netw. | 10 |
| 2026 | Transient Resource Provisioning for Connected Autonomous Vehicles-Oriented Edge Slicing: A Learning-Based Two-Timescale ApproachabstractEdge slicing is envisioned to support connected autonomous vehicle (CAV) applications with diverse key performance indicator (KPI) requirements by splitting the shared physical infrastructure into several virtual networks. Unfortunately, existing provisioning approaches struggle to accommodate the spatiotemporal dynamics of CAV traffic, leading to significant violated KPIs or soared resource usage. In this paper, we introduce the transient sharing mechanism among edge slices to obtain reused gains without generating harmful performance interference, in which a slice is allowed to access to the under-utilized reserved resources of other slices but may experience interruptions at any time. Considering the heterogeneity and uncertainty of transient resources, we further develop a two-timescale provisioning scheme. Specifically, slices proactively make reservation decisions based on multi-armed bandit architectures at the beginning of large timescales, while hinging on cost-incentive auction mechanisms selectively preempt transient resources in terms of real-time application demands at each small timescale. With extensive experiments based on real traffic traces, we demonstrate that the proposed scheme can improve 10.43% resource utilization and make slices reduce 42.92% cost than state-of-the-art works, which verifies its high assurance and adaptability. Yu Liu 0016, Jingyu Wang 0001, Qi Qi 0001, Dezhi Chen, Zirui Zhuang, Jianxin Liao, Zhu Han 0001 |
IEEE Trans. Netw. | 6 |
| 2026 | Hammurabi: Establish Cooperative Order From Pre-Trained Policies in Multi-UAV Networks
Dezhi Chen, Hongchuan He, Qi Qi 0001, Jingyu Wang 0001, Rongxin Han, Bo He 0003, Zirui Zhuang, Qianlong Fu, Jianxin Liao, Zhu Han 0001 |
IEEE Trans. Parallel Distributed Syst. | 9 |
| 2026 | Enhancing MLLMs for Online Understanding in Video Services via Preference OptimizationabstractOnline video understanding is pivotal for emerging video streaming services. However, existing Multimodal Large Language Models (MLLMs) encounter significant challenges in this domain, specifically in maintaining holistic visual perception under strict token budgets and attaining grounded semantic reasoning amidst dynamic context changes. To mitigate these issues, we propose OV-DPO, a systematic multimodal preference optimization approach. To ensure holistic visual perception, we introduce a visual preference objective that compels the model to strictly ground its reasoning in clear visual evidence. This is supported by the QGDFR strategy, which dynamically reallocates resolution budgets to preserve critical visual details for chosen samples while intentionally degrading them for rejected ones to construct valid visual contrasts. Concurrently, to promote grounded semantic reasoning, we employ a textual preference objective. We design the SROVA framework to construct high-quality preference pairs, utilizing self-refinement to generate factual chosen responses and visual degradation to induce language-prior-driven hallucinations as rejected samples. By jointly optimizing these objectives alongside anchored and supervised fine-tuning terms, OV-DPO ensures training stability and robust alignment. Experimental results demonstrate that our approach significantly outperforms baselines such as SFT and DPO. Notably, utilizing only 29K high-quality samples, our 3B-scale model achieves results competitive with significantly larger-scale models on both online and offline benchmarks, offering an efficient solution for online video understanding. Qi Qi 0001, Yixiao He, Menghao Zhang 0004, Haifeng Sun 0001, Pengfei Ren 0001, Huazheng Wang, Jianxin Liao, Jingyu Wang 0001 |
IEEE Trans. Serv. Comput. | 7 |
| 2025 | ChatTime: A Unified Multimodal Time Series Foundation Model Bridging Numerical and Textual DataabstractHuman experts typically integrate numerical and textual multimodal information to analyze time series. However, most traditional deep learning predictors rely solely on unimodal numerical data, using a fixed-length window for training and prediction on a single dataset, and cannot adapt to different scenarios. The powered pre-trained large language model has introduced new opportunities for time series analysis. Yet, existing methods are either inefficient in training, incapable of handling textual information, or lack zero-shot forecasting capability. In this paper, we innovatively model time series as a foreign language and construct ChatTime, a unified framework for time series and text processing. As an out-of-the-box multimodal time series foundation model, ChatTime provides zero-shot forecasting capability and supports bimodal input/output for both time series and text. We design a series of experiments to verify the superior performance of ChatTime across multiple tasks and scenarios, and create four multimodal datasets to address data gaps. The experimental results demonstrate the potential and utility of ChatTime. Chengsen Wang, Qi Qi 0001, Jingyu Wang 0001, Haifeng Sun 0001, Zirui Zhuang, Lei Zhang 0094, Jianxin Liao |
AAAI | 8 |
| 2025 | ClusterAttn: KV Cache Compression under Intrinsic Attention ClusteringabstractSparse attention can effectively alleviate the significant demands on memory when large language models (LLMs) process long contexts. Existing methods typically apply the same sparse pattern across different attention heads and inputs. However, this uniform approach fails to capture the inherent diversity of attention patterns within LLMs — the intrinsic attention clustering. To address this, we propose ClusterAttn, a training-free sparse attention method that provides an efficient prompt cache compression scheme under intrinsic attention clustering for efficient LLM inference.Our findings show that attention heads consistently focus on specific clusters of the prompt during decoding, a pattern detectable from an observation window at the prompt’s end. ClusterAttn adaptively fits these clusters utilizing a density-based attention clustering algorithm, thus compressing the KV cache of the prompt. Evaluations on different models across various benchmarks demonstrate ClusterAttn’s superior compression rates and efficiency. By utilizing only 1024 tokens, it can reduce memory usage by 10%–65%, resulting in a latency reduction of 12%–23% and a throughput increase of 2.6–4.8 times, all with nearly no accuracy loss. Additionally, ClusterAttn can handle up to 128k context on a single A100-80GB GPU, outperforming existing methods. Minwei Zhang, Haifeng Sun 0001, Jingyu Wang 0001, Shaolong Li, Wanyi Ning, Qi Qi 0001, Zirui Zhuang, Jianxin Liao |
ACL (1) | 8 |
| 2025 | Pose-Guided Temporal Enhancement for Robust Low-Resolution Hand Reconstructionabstract3D hand reconstruction is essential in non-contact human-computer interaction applications, but existing methods struggle with low-resolution images, which occur in slightly distant interactive scenes. Leveraging temporal information can mitigate the limitations of individual low-resolution images that lack detailed appearance information, enhancing the robustness and accuracy of hand reconstruction. Existing temporal methods typically use joint features to represent temporal information, avoiding interference from redundant background information. However, joint features excessively disregard the spatial context of visual features, limiting hand reconstruction accuracy. We propose to integrate temporal joint features with visual features to construct a robust low-resolution visual representation. We adopt Triplane Features, a dense representation with 3D spatial awareness, to bridge the gap between the joint features and visual features that are misaligned in terms of representation form and semantics. Triplane Features are obtained by orthogonally projecting joint features, embedding hand structure information into the 3D spatial context. Furthermore, we compress the spatial information of the three planes into a 2D dense feature thourgh Spatial-Aware Fusion to enhance the visual features. By using enhanced visual features enriched with temporal information for hand reconstruction, our method achieves competitive performance at much lower resolutions compared to state-of-the-art methods operating at high resolution on DexYCB, HanCo and H2O. Code is available at https://github.com/NewbieFan/Temp-LowRes-hand. Kaixin Fan, Pengfei Ren 0001, Jingyu Wang 0001, Haifeng Sun 0001, Qi Qi 0001, Zirui Zhuang, Jianxin Liao |
CVPR | 7 |
| 2025 | From Static to Dynamic: GNNs-Driven Clinical Decision-Making Assistance
Zirui Zhuang, Qi Qi 0001, Jingyu Wang 0001, Jianxin Liao, Jiachang Hao, Haifeng Sun 0001 |
DASFAA (2) | 5 |
| 2025 | Unveiling Internal Reasoning Modes in LLMs: A Deep Dive into Latent Reasoning vs. Factual Shortcuts with Attribute Rate RatioabstractYiran Yang, Haifeng Sun, Jingyu Wang, Qi Qi, Zirui Zhuang, Huazheng Wang, Pengfei Ren, Jing Wang, Jianxin Liao. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Haifeng Sun 0001, Jingyu Wang 0001, Qi Qi 0001, Zirui Zhuang, Huazheng Wang, Pengfei Ren 0001, Jing Wang 0039, Jianxin Liao |
EMNLP | 9 |
| 2025 | Atlas: Towards Real-Time Verification in Large-Scale Networks via a Native Distributed ArchitectureabstractData plane verification (DPV) can be critical in ensuring the network operates correctly. To be useful in practice, they need to be: (1) fast so as to prevent significant packet loss or security violations; (2) scalable so as to accommodate today's large-scale network architecture. Current DPV tools struggle to meet these requirements due to their centralized architecture. To be concrete, there is a bottleneck for a single-point server to perform real-time DPV tasks. Furthermore, a single-point server makes it hard to collect real-time data plane updates from every device in large-scale networks. Jingyu Wang 0001, Bo He 0003, Chenyang Zhao 0005, Qi Qi 0001, Zirui Zhuang, Haifeng Sun 0001, Lingqi Guo, Yuebin Guo, Gong Zhang 0001, Jianxin Liao |
EuroSys | 12 |
| 2025 | Prior-Aware Dynamic Temporal Modeling Framework for Sequential 3D Hand Pose Estimation
Pengfei Ren 0001, Jingyu Wang 0001, Haifeng Sun 0001, Qi Qi 0001, Menghao Zhang 0004, Lei Zhang 0094, Jing Wang 0039, Jianxin Liao |
ICCV | 9 |
| 2025 | A³-Net: Calibration-Free Multi-View 3D Hand Reconstruction for Enhanced Musical Instrument LearningabstractPrecise 3D hand posture is essential for learning musical instruments. Reconstructing highly precise 3D hand gestures enables learners to correct and master proper techniques through 3D simulation and Extended Reality. However, exsiting methods typically rely on precisely calibrated multi-camera systems, which are not easily deployable in everyday environments. In this paper, we focus on calibration-free multi-view 3D hand reconstruction in unconstrained scenarios. Establishing correspondences between multi-view images is particularly challenging without camera extrinsics. To address this, we propose A^3-Net, a multi-level alignment framework that utilizes 3D structural representations with hierarchical geometric and explicit semantic information as alignment proxies, facilitating multi-view feature interaction in both 3D geometric space and 2D visual space. Specifically, we first perfrom global geometric alignment to map multi-view features into a canonical space. Subsequently, we aggregate information into predefined sparse and dense proxies to further integrate cross-view semantics through mutual interaction. Finnaly, we perfrom 2D alignment to align projected 2D visual features with 2D observations. Our method achieves state-of-the-art results in the multi-view 3D hand reconstruction task, demonstrating the effectiveness of our proposed framework. Geng Chen 0006, Xufeng Jian, Pengfei Ren 0001, Jingyu Wang 0001, Haifeng Sun 0001, Qi Qi 0001, Jing Wang 0039, Jianxin Liao |
IJCAI | 9 |
| 2025 | Efficient Inter-Operator Scheduling for Concurrent Recommendation Model Inference on GPUabstractDeep learning-based recommendation systems are increasingly important in the industry. To meet strict SLA requirements, serving frameworks must efficiently handle concurrent queries. However, current serving systems fail to serve concurrent queries due to the following problems: (1) inefficient operator (op) scheduling due to the query-wise op launching mechanism, and (2) heavy contention caused by the mutable nature of recommendation model inference. This paper presents RecOS, a system designed to optimize concurrent recommendation model inference on GPUs. RecOS efficiently schedules ops from different queries by monitoring GPU workloads and assigning ops to the most suitable streams. This approach reduces contention and enhances inference efficiency by leveraging inter-op parallelism and op characteristics. To maintain correctness across multiple CUDA streams, RecOS introduces a unified asynchronous tensor management mechanism. Evaluations demonstrate that RecOS improves online service performance, reducing latency by up to 68%. Shuxi Guo, Qi Qi 0001, Haifeng Sun 0001, Jianxin Liao, Jingyu Wang 0001 |
IJCAI | 8 |
| 2025 | Beyond Statistical Analysis: Multimodal Framework for Time Series Forecasting with LLM-Driven Temporal PatternabstractAccurate forecasting of time series is crucial for many applications in the real world. Conventional methods primarily rely on statistical analysis of historical data, often leading to overfitting and failing to account for background information and constraints imposed by external events. Therefore, introducing large language models (LLMs) with robust textual capabilities holds significant potential. However, due to the inherent limitations of LLMs in handling numerical data, they do not exhibit advantages in precise numerical prediction tasks. Therefore, we propose a framework to integrate LLMs with conventional methods synergistically. Rather than directly outputting numerical predictions, we leverage the capabilities of the LLMs to generate textual temporal patterns, thereby fully utilizing their inherent knowledge and reasoning abilities. Additionally, we introduce a memory network designed to decode these textual representations into a format that numerical models can effectively interpret. This approach not only capitalizes on the strengths of the LLM in text processing but also bridges the gap between textual and numerical data, enhancing the overall predictive performance of the model. Our experimental results demonstrate the framework's effectiveness, achieving state-of-the-art performance on various benchmark datasets. Jiahong Xiong, Chengsen Wang, Haifeng Sun 0001, Yuhan Jing, Qi Qi 0001, Zirui Zhuang, Lei Zhang 0094, Jianxin Liao, Jingyu Wang 0001 |
IJCAI | 8 |
| 2025 | Network CoPilot: Intent-Driven Network Configuration Updating for Service Guarantee
Rongxin Han, Jingyu Wang 0001, Haifeng Sun 0001, Zengteng Jiang, Qi Qi 0001, Zirui Zhuang, Jianxin Liao |
INFOCOM | 8 |
| 2025 | Rule Meets Learning: Confidence-Aware Multi-View Fusion for Self-Supervised 3D Hand Pose EstimationabstractSelf-supervised 3D hand pose estimation methods can leverage labeled synthetic data along with unlabeled real-world data for model training, thereby alleviating the reliance on large-scale annotated datasets. Multi-view information fusion is a key factor in the success of these methods. Rule-based fixed fusion methods are simple, efficient, and generalizable, but they neglect the rich visual information in each view. Neural network-based learnable fusion methods can effectively model both intra- and inter-view semantic context, but they tend to overfit to the domain-specific feature of synthetic data and susceptible to interference of domain gaps. In this paper, we decompose multi-view fusion into two components: a learnable confidence estimation stage and a fixed confidence fusion stage. This design not only enables effective use of multi-view semantic cues but also ensures strong cross-domain generalization. To achieve accurate and robust confidence estimation, our method jointly exploits both multi-view pose consistency and pose-to-data consistency. Experiments on three public datasets demonstrate that our approach significantly outperforms existing state-of-the-art self-supervised 3D hand pose estimation methods. Pengfei Ren 0001, Jingyu Wang 0001, Haifeng Sun 0001, Qi Qi 0001, Jing Wang 0039, Jianxin Liao |
ACM Multimedia | 6 |
| 2025 | A Dual-Branch 3D Spatial-Aware Latent Diffusion for Realistic Depth Image SynthesisabstractSynthetic images serve as a promising alternative to real images in 3D hand pose estimation, providing accurate annotations at a lower cost. However, the domain gap between real and synthetic images constrains the generalization ability of hand pose estimation trained on synthetic data. Previous methods rely on Generative Adversarial Networks (GANs) for domain translation; however, they fail to achieve realistic depth synthesis due to instability and limited image quality. Diffusion models provide high-quality synthesis due to their stability and controllability. However, existing methods often ignore the 3D structure awareness in hand image generation. In this paper, we propose a Dual-Branch 3D Spatial-Aware Latent Diffusion (DSW-LD) for realistic depth image generation. The Global Structure Module (GSM) and the Local Geometry Module (LGM) complement each other, with GSM capturing global spatial structure through coarse-grained 3D joint features and LGM focusing on local geometric details using fine-grained 3D mesh representations. To maintain the global structure consistency, we adopt a layer-aware injection mechanism that enables the model to adaptively learn the optimal representation from fused 2D latent representations and 3D joint features. To explicitly align 3D and 2D features of local regions and enhance the flexibility of feature matching, we design a dynamic depth-aware interpolation to project 3D mesh features into 2D image space. Both quantitative and qualitative experimental results demonstrate the superiority of our method over the state-of-the-arts for realistic depth synthesis. Compared to training only on real depth images, our method enables the hand pose estimator to achieve significantly better performance with our synthetic data and less real data (10%). Shuang Hao 0017, Pengfei Ren 0001, Lei Zhang 0094, Haifeng Sun 0001, Pan Ting, Menghao Zhang 0004, Cong Liu 0046, Qi Qi 0001, Jianxin Liao, Jingyu Wang 0001 |
ACM Multimedia | 9 |
| 2025 | Foresail: LLM Sensor Knowledge Empowered Status-guided Network for Multivariate Time-series ClassificationabstractMultivariate time-series (MTS) classification tasks play a key role in data-driven applications spanning healthcare, finance, and mobile communication. As MTS data are typically collected from multiple interdependent sensors, the resulting temporal patterns inherently reflect the characteristics of the underlying sensing systems. Despite this connection, conventional MTS classification models predominantly focus on raw time-series data while disregarding valuable sensor-specific prior knowledge, which fundamentally constrains their classification accuracy. The emergence of large language models (LLMs) has encoded extensive sensor-related knowledge within their parameter spaces. However, effectively harnessing such knowledge to enhance MTS classification networks remains an open challenge. To address this, we propose Foresail, a status-guided neural framework that bridges this gap through systematic integration of LLM-derived sensor knowledge via the status relationship matrix and fine-grained status labels. Foresail can be seamlessly integrated with existing MTS networks to optimize performance and generate interpretable intermediate results. Experiments on irregularly and regularly sampled MTS data demonstrate that Foresail outperforms state-of-the-art approaches, achieving a notable improvement in F1-score of up to 10.9% compared to the basic MTS network. Yuhan Jing, Bo He 0003, Haifeng Sun 0001, Qi Qi 0001, Zirui Zhuang, Lei Zhang 0094, Jianxin Liao, Jingyu Wang 0001 |
ACM Multimedia | 7 |
| 2025 | Watch, Skip, Repeat: Hotspot-Aware Joint Optimization for Video StreamingabstractVideo streaming platforms and existing ABRs traditionally assume uninterrupted sequential playback, yet users frequently skip to points of interest-a fundamental mismatch causing degradation of quality of experience at high-interest segments while wasting bandwidth on skipped content. We address this through our Hotspot-Aware Joint Optimization framework, which reframes video streaming as a non-monotonic optimization problem with discontinuous state transitions caused by navigation events. Our framework jointly optimizes adaptive bitrate decisions and buffer management by leveraging viewer engagement patterns to predict navigation behavior. Our approach combines: (1) a mathematical formulation capturing state discontinuities in non-sequential viewing, (2) self-supervised models predicting navigation targets using only aggregate viewing data, and (3) hotspot-aware ABR and buffer management algorithms implemented through our Streaming Local Search (SLS) technique that dynamically prioritize quality for frequently-watched segments. Evaluation across diverse content and network conditions demonstrates our framework delivers 38.2% higher quality in hotspot regions, 32.5% reduced navigation delays, and 27.1% improved resource efficiency compared to traditional methods. These improvements establish a foundation for streaming systems that adapt to both network conditions and content structure, aligning resource allocation with actual viewing patterns. Daoxu Sheng, Qi Qi 0001, Jingyu Wang 0001, Jianxin Liao |
ACM Multimedia | 4 |
| 2025 | Unified 2D-3D Discrete Priors for Noise-Robust and Calibration-Free Multiview 3D Human Pose EstimationabstractMulti-view 3D human pose estimation (HPE) leverages complementary information across views to improve accuracy and robustness. Traditional methods rely on camera calibration to establish geometric correspondences, which is sensitive to calibration accuracy and lacks flexibility in dynamic settings. Calibration-free approaches address these limitations by learning adaptive view interactions, typically leveraging expressive and flexible continuous representations. However, as the multiview interaction relationship is learned entirely from data without constraint, they are vulnerable to noisy input, which can propagate, amplify and accumulate errors across all views, severely corrupting the final estimated pose.
To mitigate this, we propose a novel framework that integrates a noise-resilient discrete prior into the continuous representation-based model. Specifically, we introduce the \textit{UniCodebook}, a unified, compact, robust, and discrete representation complementary to continuous features, allowing the model to benefit from robustness to noise while preserving regression capability.
Furthermore, we further propose an attribute-preserving and complementarity-enhancing Discrete-Continuous Spatial Attention (DCSA) mechanism to facilitate interaction between discrete priors and continuous pose features.
Extensive experiments on three representative datasets demonstrate that our approach outperforms both calibration-required and calibration-free methods, achieving state-of-the-art performance. Geng Chen 0006, Pengfei Ren 0001, Xufeng Jian, Haifeng Sun 0001, Menghao Zhang 0004, Qi Qi 0001, Zirui Zhuang, Jing Wang 0039, Jianxin Liao, Jingyu Wang 0001 |
NeurIPS | 9 |
| 2025 | Generalizable Hand-Object Modeling from Monocular RGB Images via 3D GaussiansabstractRecent advances in hand-object interaction modeling have employed implicit representations, such as Signed Distance Functions (SDF) and Neural Radiance Fields (NeRF) to reconstruct hands and objects with arbitrary topology and photo-realistic detail. However, these methods often rely on dense 3D surface annotations, or are tailored to short clips constrained in motion trajectories and scene contexts, limiting their generalization to diverse environments and movement patterns. In this work, we present HOGS, an adaptively perceptive 3D Gaussian Splatting (3DGS) framework for generalizable hand-object modeling from unconstrained monocular RGB images. By integrating photometric cues from the visual modality with the physically grounded structure of 3D Gaussians, HOGS disentangles inherent geometry from transient lighting and motion-induced appearance changes. This endows hand-object assets with the ability to generalize to unseen environments and dynamic motion patterns. Experiments on two challenging datasets demonstrate that HOGS outperforms state-of-the-art methods in monocular hand-object reconstruction and photo-realistic rendering. Pengfei Ren 0001, Qi Qi 0001, Haifeng Sun 0001, Zirui Zhuang, Jing Wang 0039, Jianxin Liao, Jingyu Wang 0001 |
NeurIPS | 7 |
| 2025 | Do LVLMs Truly Understand Video Anomalies? Revealing Hallucination via Co-Occurrence PatternsabstractLarge Vision-Language Models (LVLMs) pretrained on large-scale multimodal data have shown promising capabilities in Video Anomaly Detection (VAD). However, their ability to reason about abnormal events based on scene semantics remains underexplored. In this paper, we investigate LVLMs’ behavior in VAD from a visual-textual co-occurrence perspective, focusing on whether their decisions are driven by statistical shortcuts between visual instances and textual phrases. By analyzing visual-textual co-occurrence in pretraining data and conducting experiments under different data settings, we reveal a hallucination phenomenon: LVLMs tend to rely on co-occurrence patterns between visual instances and textual phrases associated with either normality or abnormality, leading to incorrect predictions when these high-frequency objects appear in semantically mismatched contexts. To address this issue, we propose VAD-DPO, a direct preference optimization method supervised with counter-example pairs. By constructing visually similar but semantically contrasting video clips, VAD-DPO encourages the model to align its predictions with the semantics of scene rather than relying on co-occurrence patterns. Extensive experiments on six benchmark datasets demonstrate the effectiveness of VAD-DPO in enhancing both anomaly detection and reasoning performance, particularly in scene-dependent scenarios. Menghao Zhang 0004, Huazheng Wang, Pengfei Ren 0001, Kangheng Lin, Qi Qi 0001, Haifeng Sun 0001, Zirui Zhuang, Lei Zhang 0094, Jianxin Liao, Jingyu Wang 0001 |
NeurIPS | 9 |
| 2025 | RACE: Operator Choreography for Inference Acceleration in Personalized Recommender SystemabstractWeb recommender systems play a crucial role in enhancing user experience and engagement by providing personalized content suggestions. These systems often leverage Recommendation Models (RMs) that integrate dense Deep Neural Networks (DNNs) with traditional sparse components to improve accuracy. The inference process for RMs in web recommender systems is typically executed in a CPU-GPU hybrid environment to handle the computational complexity. However, existing approaches have struggled to fully optimize RM inference and achieve real-time recommendations due to the challenges posed by the intricate RM operator graphs and heterogeneous hardware architecture. This paper presents RACE, an operator choreography framework designed to accelerate RM inference specifically for web recommender systems. RACE introduces three key components: an Operator Graph Placement (OGP) module for strategic operator placement, an Inference Embedding Table (IET) for efficient GPU lookups, and an Automatic Subgraph Fusion (ASF) module to optimize GPU utilization. Evaluated in production settings, RACE achieves a 3.2× speedup in RM inference, significantly enhancing the performance and responsiveness of web recommender systems. Shaolong Li, Jiaxing Liu 0006, Qi Qi 0001, Jianxin Liao, Jingyu Wang 0001 |
SPAA | 5 |
| 2025 | NetKeeper: Enhancing Network Resilience with Autonomous Network Configuration Update on Traffic Patterns and Anomalies
Zhaoyang Wan, Rongxin Han, Haifeng Sun 0001, Qi Qi 0001, Zirui Zhuang, Bo He 0003, Jianxin Liao, Jingyu Wang 0001 |
USENIX ATC | 8 |
| 2025 | Robustness Verification of Deep Graph Neural Networks Tightened by Linear ApproximationabstractRecent research indicates that adding residual connections in Graph Neural Networks (GNNs) would amplify susceptibility to anomalous nodes, consequently undermining the robustness of deep GNNs in practical settings. However, existing verification methods encounter challenges with the increasing number of parameters and computational overhead in deep GNNs. In this paper, we derive the general form of the residual connections and apply the dual backpropagation network to deep GNNs. Considering the heightened computational errors arising from the increased number of layers in deep GNNs, we propose a new method for calculating intermediate activation bounds of GNNs based on linear approximation. Experimental results show that new method can effectively enhance the verification accuracy. Notably, the maximum perturbation value of nodes correctly classified shows an average improvement of 119.5%. To showcase the the efficacy and scalability of our method, we verify robustness of deep GNNs on six different graph datasets, and our method can effectively verify the robustness of deep GNNs even with 32 layers of residual connections, i.e. verify over 87.29% of nodes in the Citeseer dataset. Furthermore, we analyse the influence of the graph structural properties on the robustness of the model. Xingyu Zeng, Qi Qi 0001, Jingyu Wang 0001, Haodong Deng, Haifeng Sun 0001, Zirui Zhuang, Jianxin Liao |
WSDM | 8 |
| 2025 | Cache-INT: In-network caching-enabled In-band Network Telemetry
Hua Zhang 0002, Yuqi Dai, Yibo Pi, Jingyu Wang 0001, Jianxin Liao |
Comput. Networks | 6 |
| 2025 | INT-LLPP: Lightweight in-band network-wide telemetry with low-latency and low-overhead path planning
Hua Zhang 0002, Yuqi Dai, Cheng Zeng 0002, Jingyu Wang 0001, Jianxin Liao |
Comput. Commun. | 6 |
| 2025 | NTP-INT: Network traffic prediction-driven in-band network telemetry for high-load switches
Hua Zhang 0002, Yuqi Dai, Cheng Zeng 0002, Jingyu Wang 0001, Jianxin Liao |
J. Netw. Comput. Appl. | 6 |
| 2025 | Towards Bare-Hand Interaction for Whiteboard Collaboration in Virtual RealityabstractWhiteboard collaboration in virtual reality (VR) is an important task in collaborative virtual environments. The current research mainly relies on the use of controllers or dedicated pens but additional devices will cause inconvenience to users. Bare-hand writing offers rich collaborative semantics through natural gestures but remains underexplored. This paper addresses challenges and solutions for bare-hand whiteboard collaboration. We analyze the input process and identify key challenges in determining pen-drop, writing, and pen-lift intentions while maintaining user control over their avatar. Our approach addresses two VR scenarios: one without and one with physical planes. The method for the first case is called Air-writing, which dynamically adjusts the distance between the avatar's torso and the virtual whiteboard during the processes of pen-drop and pen-lift to ensure a consistent writing experience in VR. The method for the second case is called Physical-writing, which allows users to write smoothly with passive haptic feedback and physical constraints provided by the real surface by remapping the whiteboard in VR with a plane in reality. A comprehensive user study is conducted to evaluate communication efficiency, input accuracy, collaboration efficiency, and user experience of the two methods. The experimental results indicate that bare-hand interaction improves communication efficiency by 8% over controllers and performs similarly to real-world whiteboard collaboration. The Physical-writing method also demonstrates higher accuracy and user satisfaction compared to the Air-writing method. Guangtian Liu, Haonan Su, Jingyu Wang 0001, Qi Qi 0001, Haifeng Sun 0001, Zirui Zhuang, Pengfei Ren 0001, Jianxin Liao |
Proc. ACM Hum. Comput. Interact. | 8 |
| 2025 | DeepZoning: Re-accelerate CNN Inference with Zoning Graph for Heterogeneous Edge ClusterabstractParallelizing CNN inference on heterogeneous edge clusters with data parallelism has gained popularity as a way to meet real-time requirements without sacrificing model accuracy. However, existing algorithms struggle to find optimal parallel granularity for complex CNNS, the structure of which is a directed acyclic graph (DAG) rather than a chain, and the parallel dimension is inflexible. To distribute the workload of modern CNNs on heterogeneous devices is also proven as NP-hard problem. In this article, we introduce DeepZoning , a versatile and cooperative inference framework that combines both model and data parallelism to accelerate CNN inference. DeepZoning employs two algorithms at different levels: (1) a low-level Adaptive Workload Partition algorithm that uses linear programming and takes spatial and channel dimensions into optimization during the search for feature map distribution on heterogeneous devices, and (2) a high-level Model Partition algorithm that finds the optimal model granularity and organizes complex CNNs into sequential zones to balance communication and computation during execution. Our experimental evaluations show that DeepZoning is effective, achieving up to a 3.02× speed improvement on our experimental prototype compared to state-of-the-art algorithms. Jingyu Wang 0001, Ruilong Ma, Qi Qi 0001, Zirui Zhuang, Jing Wang 0039, Jianxin Liao, Song Guo 0001 |
ACM Trans. Archit. Code Optim. | 7 |
| 2025 | Intent-Based Autonomous Network Framework Guided by Large Language ModelabstractWith the rapid development of next-generation networks, the highly heterogeneous and dynamic nature of networks poses significant challenges for automated network management. Autonomous Network (AN), as a new network paradigm, aims to provide customers with a zero-wait, zero-touch, and zero-fault experience. AN facilitates network management through intent-driven interactions and provides on-demand resource orchestration and service scheduling. However, accurately translating user intents into commands and allocating resources on demand for services remain significant challenges for AN. Therefore, this paper proposes IAN, an intent-based AN framework guided by the Large Language Model (LLM). In the intent translation phase, IAN introduces RAG to enhance command generation quality by retrieving from manuals. In the resource allocation phase, the method utilizes LLM to analyze service characteristics, thereby guiding the training and inference of the resource allocation model to effectively distribute resources uniformly across emerging services. Experimental results demonstrate that IAN improves performance by 52.66% in intent translation tasks and increases overall gain by 33.57% in resource allocation tasks compared to other models. Lingqi Guo, Lei Zhang 0094, Jingyu Wang 0001, Haifeng Sun 0001, Bo He 0003, Qi Qi 0001, Jianxin Liao |
IEEE Trans Autom. Sci. Eng. | 9 |
| 2025 | Flight Trajectory Control With Network-Oriented Hierarchical Reinforcement Learning for UAVs-Assisted Data Time-Sensitive IoTabstractWithin the Internet of Things (IoT) for traffic monitoring, the employment of autonomous aerial vehicles (AAVs) as relays for collecting and transmitting real-time data from traffic sensors to base stations has proven a promising approach. In UAV-assisted Data Time-Sensitive IoT (DTIoT), the Age of Information is a crucial metric assessing data freshness, measuring the elapsed time from traffic sensors to the base station. Optimizing flight trajectories of multiple UAVs to minimize AoI while adhering to energy constraints poses a significant challenge. Current research often employs deep reinforcement learning for UAV trajectory control. Nevertheless, managing multi-agent continuous trajectories in intricate DTIoT network environments faces obstacles due to sparse rewards, thus impeding the training of deep neural network-based control policies using traditional DRL techniques. In this paper, we propose a network-oriented hierarchical reinforcement learning (NO-HRL) to control the UAVs’ flight trajectory in DTIoT networks for minimizing the AoI. We devise a control policy leveraging a two-tier hierarchical DRL framework, with the upper tier determining the target and the lower tier executing it. We also introduce a decoupled sequential training approach to efficiently train the mutually dependent two-tier DRL network of NO-HRL. Experimental results demonstrate that our method excels in optimizing AoI for DTIoT compared to other baselines. Jingyu Wang 0001, Dezhi Chen, Qianlong Fu, Qi Qi 0001, Haifeng Sun 0001, Bo He 0003, Jianxin Liao |
IEEE Trans. Intell. Transp. Syst. | 8 |
| 2025 | Hierarchical Index Retrieval-Driven Wireless Network Intent Translation With LLMabstractIntent-Based Networking (IBN) represents an emerging network management concept that is designed to fulfill user service requirements through automation. At its core, IBN is capable of translating user intent into network policies, thereby enabling automated configuration and management. However, the application of IBN has been limited by challenges associated with automation and intelligence. The recent widespread adoption of Large Language Model (LLM) has partially mitigated these issues. Nonetheless, hardware heterogeneity and high dynamic networks remain significant challenges for IBN: (i) Devices from different vendors are challenging to manage uniformly; (ii) Aligning service demands with rapidly changing network status is difficult. To address these challenges, we propose LIT, a framework of LLM-empowered Intent Translation with manual guidance. LIT incorporates Retrieval-Augmented Generation (RAG) to reference hardware manuals and enhance the generation results of LLMs. To reduce noise from retrieval results, we optimized the general RAG process. Additionally, LIT introduces MoE (Mixture of Experts) to adjust parameter values according to network status by synthesizing results from multiple expert models. Experiments demonstrate that LIT alleviates the challenges faced by IBN, achieving a 57.5% improvement in F1 score compared to the baseline. Jingyu Wang 0001, Lingqi Guo, Caijun Yan, Haifeng Sun 0001, Lei Zhang 0094, Zirui Zhuang, Qi Qi 0001, Jianxin Liao |
IEEE Trans. Mob. Comput. | 9 |
| 2025 | Federated Fine-Tuning on Heterogeneous LoRAs With Error-Compensated AggregationabstractFederated learning (FL) has recently been applied to the parameter-efficient fine-tuning (PEFT) of large language models (LLMs). While promising, client resource heterogeneity has imposed the challenge of the "bucket effect" to FL, where model configuration must cater to the client with the fewest resources. To tackle this issue, heterogeneous low-rank adaptation (LoRA) has recently emerged in FL, which enables clients to do local fine-tuning with different LoRA ranks. However, existing works in this area typically adopt zero-padding, stacking, or singular value decomposition (SVD) for LoRA aggregation, which often incur precision loss or significant overhead, limiting their practicality. In this article, we propose ECLoRA, a novel method for federated fine-tuning with heterogeneous LoRA settings across clients. ECLoRA employs randomized SVD (RSVD) to dramatically reduce aggregation overhead while introducing an error compensation (EC) mechanism that incorporates the decomposition error from previous rounds to improve aggregation precision. Extensive experiments on four widely used foundation models across six public tasks demonstrate the effectiveness of ECLoRA. Specifically, ECLoRA is: (1) accurate, significantly improving the final model performance; (2) fast, accelerating convergence with an average speedup of $1.54\times $ to $3.01\times $ ; and (3) practical, reducing aggregation time by approximately $40\times $ compared to classical SVD. Wanyi Ning, Jingyu Wang 0001, Qi Qi 0001, Haifeng Sun 0001, Daixuan Cheng, Cong Liu 0046, Lei Zhang 0094, Zirui Zhuang, Jianxin Liao |
IEEE Trans. Neural Networks Learn. Syst. | 9 |
| 2025 | Fast and Scalable Data Plane Verification for Burst Updates With Edge-PredicateabstractThere is an increasing interest in data plane verification, which is designed to automatically verify network correctness through directly analyzing the data plane. Recent data plane verifiers have been able to do real-time sub-millisecond per rule verification. However, we observe that in real-world networks, individual data plane updates rarely occur. On the contrary, there are always a certain number of updates generated in a short period of time, called asburst updates, due to high-level user intend or uncertain network events. When it comes to this real-life scenario, yet, the current equivalence class (EC) based methods are unable to solve themodel-wide changesproblem caused by the EC itself, which significantly slows down the verification speed of burst updates. To overcome this limitation, we present EPVerifier, a fast, scalable data plane verifier accelerating burst updates verification with edge-predicate (EP). Instead of classifying packets into ECs according to global forwarding behavior, the EPVerifier uses one EP per edge to represent all packets that can pass through. Furthermore, with EPs that clearly have localized properties, we introduce a rule type extension that does not require a change in the granularity of the network model to support ACLs and NATs that are prevalent in real devices, and obtain better-performing parallelism by dividing the verification task based on switches. Experiments on both dataset simulations and real-life deployments show that EPVerifier achieves 2-$10\times $faster data plane verification than the state-of-the-art and such advantage expand with the data plane’s complexity and update scale growth. Jingyu Wang 0001, Chenyang Zhao 0005, Zirui Zhuang, Qi Qi 0001, Yuebin Guo, Haifeng Sun 0001, Lingqi Guo, Jianxin Liao |
IEEE Trans. Netw. | 8 |
| 2025 | LogNotion: Highlighting Massive Logs to Assist Human Reading and Decision MakingabstractMassive logs contain crucial information about the working status of software systems, which contributes to anomaly detection and troubleshooting. For engineers, it is a laborious task to manually inspect raw logs to know the system running status, and therefore an automated log summarization tool can be helpful. However, due to the specificity of logs in terms of grammar, vocabulary and semantics, existing natural language-based methods cannot perform well in log analysis. To address these issues, we propose LogNotion, a general log summarization framework that highlights the log messages to assist human reading and decision making. We first explore the role played by triplets in log analysis, and propose a triplet extraction method based on sequence tagging and component alignment, in which the specificity of logs is fully taken into account. Then, we propose an unsupervised log summarization method to extract both regular and noteworthy information based on triplets. Comprehensive experiments are conducted on seven real-world log datasets and the results show that LogNotion improves the average ROUGE-1 by 0.26, recall by 0.12, and compression ratio by 2.13%, compared to state-of-the-art log summarization tools. The helpfulness, readability and generalizability are also verified through human evaluation and cross-dataset tests. Guojun Chu, Jingyu Wang 0001, Tao Sun 0010, Qi Qi 0001, Haifeng Sun 0001, Zirui Zhuang, Jianxin Liao |
IEEE Trans. Serv. Comput. | 7 |
| 2025 | Anomaly Detection on Interleaved Log Data With Semantic Association Mining on Log-Entity GraphabstractLogs record crucial information about runtime status of software system, which can be utilized for anomaly detection and fault diagnosis. However, techniques struggle to perform effectively when dealing with interleaved logs and entities that influence each other. Although manually specifying a grouping field for each dataset can handle the single grouping scenario, the problems of multiple and heterogeneous grouping still remain unsolved. To break through these limitations, we first design a log semantic association mining approach to convert log sequences into Log-Entity Graph, and then propose a novel log anomaly detection model named Lograph. The semantic association can be utilized to implicitly group the logs and sort out complex dependencies between entities, which have been overlooked in existing literature. Also, a Heterogeneous Graph Attention Network is utilized to effectively capture anomalous patterns of both logs and entities, where Log-Entity Graph serves as a data management and feature engineering module. We evaluate our model on real-world log datasets, comparing with nine baseline models. The experimental results demonstrate that Lograph can improve the accuracy of anomaly detection, especially on the datasets where entity relationships are intricate and grouping strategies are not applicable. Guojun Chu, Jingyu Wang 0001, Qi Qi 0001, Haifeng Sun 0001, Zirui Zhuang, Bo He 0003, Yuhan Jing, Lei Zhang 0094, Jianxin Liao |
IEEE Trans. Software Eng. | 9 |
| 2024 | Keypoint Fusion for RGB-D Based 3D Hand Pose EstimationabstractPrevious 3D hand pose estimation methods primarily rely on a single modality, either RGB or depth, and the comprehensive utilization of the dual modalities has not been extensively explored. RGB and depth data provide complementary information and thus can be fused to enhance the robustness of 3D hand pose estimation. However, there exist two problems for applying existing fusion methods in 3D hand pose estimation: redundancy of dense feature fusion and ambiguity of visual features. First, pixel-wise feature interactions introduce high computational costs and ineffective calculations of invalid pixels. Second, visual features suffer from ambiguity due to color and texture similarities, as well as depth holes and noise caused by frequent hand movements, which interferes with modeling cross-modal correlations. In this paper, we propose Keypoint-Fusion for RGB-D based 3D hand pose estimation, which leverages the unique advantages of dual modalities to mutually eliminate the feature ambiguity, and performs cross-modal feature fusion in a more efficient way. Specifically, we focus cross-modal fusion on sparse yet informative spatial regions (i.e. keypoints). Meanwhile, by explicitly extracting relatively more reliable information as disambiguation evidence, depth modality provides 3D geometric information for RGB feature pixels, and RGB modality complements the precise edge information lost due to the depth noise. Keypoint-Fusion achieves state-of-the-art performance on two challenging hand datasets, significantly decreasing the error compared with previous single-modal methods. Pengfei Ren 0001, Jingyu Wang 0001, Haifeng Sun 0001, Qi Qi 0001, Zirui Zhuang, Jianxin Liao |
AAAI | 8 |
| 2024 | NetRen: Service Migration-Driven Network Renascence with Synthesizing Updated ConfigurationabstractChanges in enterprise networks require updated configurations. However, manual configurations with slow update efficiency, poor performance, and handling limitations, lead to the unavailability of updated networks. Therefore, we propose an efficient network renascence framework, NetRen, which synthesizes OSPF/BGP configurations driven by service and traffic migration. We follow the workflow of sketch extraction, configuration synthesis, and repair. Initially, comprehensive graphs are constructed to represent configuration sketches. We propose a GraphTrans synthesizer with Transformer's benefits of long-range focus and parallel reasoning. Training samples with the optimization relationship enable the synthesizer to achieve a mapping that optimizes performance based on configurations. To overcome the satisfiability barrier, configurations from the synthesizer are input to the stepwise configuration repairer as well-initialized solutions, achieving rapid configuration repair. Experiments demonstrate that the consistency of network configurations output by the GraphTrans synthesizer averages 98%. NetRen achieves a 312.4× increase in synthesis efficiency and a 5.83% improvement in network performance. Rongxin Han, Jingyu Wang 0001, Qi Qi 0001, Haifeng Sun 0001, Chaowei Xu, Zhaoyang Wan, Zirui Zhuang, Yichuan Yu, Jianxin Liao |
ASPLOS (3) | 9 |
| 2024 | Distantly Supervised Contrastive Learning for Low-Resource Scripting Language SummarizationabstractCode summarization provides a natural language description for a given piece of code. In this work, we focus on scripting code—programming languages that interact with specific devices through commands. The low-resource nature of scripting languages makes traditional code summarization methods challenging to apply. To address this, we introduce a novel framework: distantly supervised contrastive learning for low-resource scripting language summarization. This framework leverages limited atomic commands and category constraints to enhance code representations. Extensive experiments demonstrate our method’s superiority over competitive baselines. Junzhe Liang, Haifeng Sun 0001, Zirui Zhuang, Qi Qi 0001, Jingyu Wang 0001, Jianxin Liao |
LREC/COLING | 6 |
| 2024 | Dynamic Support Information Mining for Category-Agnostic Pose EstimationabstractCategory-agnostic pose estimation (CAPE) aims to predict the pose of a query image based on few support images with pose annotations. Existing methods achieve the localization of arbitrary keypoints through similarity matching between support keypoint features and query image features. However, these methods primarily focus on mining information from the query images, neglecting the fact that support samples with keypoint annotations contain rich category-specific fine-grained semantic information and prior structural information. In this paper, we propose a Support-based Dynamic Perception Network (SDP-Net) for the robust and accurate CAPE. On the one hand, SDPNet models complex dependencies between support keypoints, constructing category-specific prior structure to guide the interaction of query keypoints. On the other hand, SDPNet extracts fine-grained semantic information from support samples, dynamically modulating the refinement process of query. Our method outperforms existing methods on MP-100 dataset by a large margin. Pengfei Ren 0001, Haifeng Sun 0001, Qi Qi 0001, Jingyu Wang 0001, Jianxin Liao |
CVPR | 6 |
| 2024 | Multi-Scale Video Anomaly Detection by Multi-Grained Spatio-Temporal Representation LearningabstractRecent progress in video anomaly detection suggests that the features of appearance and motion play crucial roles in distinguishing abnormal patterns from normal ones. However, we note that the effect of spatial scales of anomalies is ignored. The fact that many abnormal events occur in limited localized regions and severe background noise in-terferes with the learning of anomalous changes. Mean-while, most existing methods are limited by coarse-grained modeling approaches, which are inadequate for learning highly discriminative features to discriminate subtle differences between small-scale anomalies and normal patterns. To this end, this paper address multi-scale video anomaly detection by multi-grained spatiotemporal representation learning. We utilize video continuity to design three proxy tasks to perform feature learning at both coarse-grained and fine-grained levels, i.e., continuity judgment, discontinuity localization, and missing frame estimation. In particular, we formulate missing frame estimation as a contrastive learning task in feature space instead of a reconstruction task in RGB space to learn highly discriminative features. Experiments show that our proposed method outperforms state-of-the-art methods on four datasets, especially in scenes with small-scale anomalies. Menghao Zhang 0004, Jingyu Wang 0001, Qi Qi 0001, Haifeng Sun 0001, Zirui Zhuang, Pengfei Ren 0001, Ruilong Ma, Jianxin Liao |
CVPR | 8 |
| 2024 | Coarse-to-Fine Implicit Representation Learning for 3D Hand-Object Reconstruction from a Single RGB-D Image
Pengfei Ren 0001, Jingyu Wang 0001, Qi Qi 0001, Haifeng Sun 0001, Zirui Zhuang, Jianxin Liao |
ECCV (51) | 7 |
| 2024 | FLUK: Protecting Federated Learning Against Malicious Clients for Internet of Vehicles
Mengde Zhu, Wanyi Ning, Qi Qi 0001, Jingyu Wang 0001, Zirui Zhuang, Haifeng Sun 0001, Jianxin Liao |
Euro-Par (2) | 8 |
| 2024 | DTG: Learning A Dynamic Token Graph for 3D Pose Forecasting
Yangliu He, Haoge Deng, Qiwei Shen 0001, Jianxin Liao |
ICANN (3) | 4 |
| 2024 | Multi - Agent Reinforcement Learning for Backscattering Data Collection in Multi-UAV IoTabstractUsing multiple unmanned aerial vehicles (UAVs) with backscatter communication to collect data from Internet of Things (IoT) devices has emerged as a promising solution. However, many existing UAVs path planning schemes for data collection suffer from performance degradation due to their limited consideration of the full collaboration of UAVs and dynamic stochastic environments. Therefore, we propose a path planning scheme for the data collection task in multi-UAV IoT based on multi-agent reinforcement learning (MARL) to minimize the task completion time. Due to the inherent asynchronous decision making among the agents, we model the path planning problem as a macro-action decentralized partially observable Markov decision process. Furthermore, we design an action mask mechanism to enhance data efficiency, which accelerates the training speed. Simulation results show that our scheme reduces the average task completion time by 15 %. Jianxin Liao, Jiangong Zheng, Tongyu Song, Jing Ren 0002, Xiong Wang 0001, Shizhong Xu, Sheng Wang 0006 |
ICC | 1 |
| 2024 | Towards Semantic Consistency: Dirichlet Energy Driven Robust Multi-Modal Entity AlignmentabstractMulti-Modal Entity Alignment (MMEA) is a pivotal task in Multi-Modal Knowledge Graphs (MMKGs), seeking to identify identical entities by leveraging associated modal attributes. However, real-world MMKGs confront the challenges of semantic inconsistency arising from diverse and incomplete data sources. This inconsistency is predominantly caused by the absence of specific modal attributes, manifesting in two distinct forms: disparities in attribute counts or the absence of certain modalities. Current methods address these issues through attribute interpolation, but their reliance on predefined distributions introduces modality noise, compromising original semantic information. Furthermore, the absence of a generalizable theoretical principle hampers progress towards achieving semantic consistency. In this work, we propose a generalizable theoretical principle by examining semantic consistency from the perspective of Dirichlet energy. Our research reveals that, in the presence of semantic inconsistency, models tend to overfit to modality noise, leading to over-smoothing and performance oscillations or declines, particularly in scenarios with a high rate of missing modality. To overcome these challenges, we propose DESAlign, a robust method addressing the over-smoothing caused by semantic inconsistency and interpolating missing semantics using existing modalities. Specifically, we devise a training strategy for multi-modal knowledge graph learning based on our proposed principle. Then, we introduce a propagation strategy that utilizes existing features to provide interpolation solutions for missing semantic features. DESAlign outperforms existing approaches across 60 benchmark splits, encompassing both monolingual and bilingual scenarios, achieving state-of-the-art performance. Experiments on splits with high missing modal attributes demonstrate its effectiveness, providing a robust MMEA solution to semantic inconsistency in real-world MMKGs. Yuanyi Wang, Haifeng Sun 0001, Jingyu Wang 0001, Wei Tang 0013, Qi Qi 0001, Shaoling Sun, Jianxin Liao |
ICDE | 8 |
| 2024 | Interdependency Matters: Graph Alignment for Multivariate Time Series Anomaly DetectionabstractAnomaly detection in multivariate time series (MTS) is crucial for various applications in data mining and industry. Current industrial methods typically approach anomaly detection as an unsupervised learning task, aiming to identify deviations by estimating the normal distribution in noisy, label-free datasets. These methods increasingly incorporate interdependencies between channels through graph structures to enhance accuracy. However, the role of interdependencies is more critical than previously understood, as shifts in interdependencies between MTS channels from normal to anomalous data are significant. This observation suggests that anomalies could be detected by changes in these interdependency graph series. To capitalize on this insight, we introduce MADGA (MTS Anomaly Detection via Graph Alignment), which redefines anomaly detection as a graph alignment (GA) problem that explicitly utilizes interdependencies for anomaly detection. MADGA dynamically transforms subsequences into graphs to capture the evolving interdependencies, and Graph alignment is performed between these graphs, optimizing an alignment plan that minimizes cost, effectively minimizing the distance for normal data and maximizing it for anomalous data. Uniquely, our GA approach involves explicit alignment of both nodes and edges, employing Wasserstein distance for nodes and Gromov-Wasserstein distance for edges. To our knowledge, this is the first application of GA to MTS anomaly detection that explicitly leverages interdependency for this purpose. Extensive experiments on diverse real-world datasets validate the effectiveness of MADGA, demonstrating its capability to detect anomalies and differentiate interdependencies, consistently achieving state-of-the-art across various scenarios. Yuanyi Wang, Haifeng Sun 0001, Chengsen Wang, Mengde Zhu, Jingyu Wang 0001, Wei Tang 0013, Qi Qi 0001, Zirui Zhuang, Jianxin Liao |
ICDM | 9 |
| 2024 | MONR: Multi-Objective Optimizing Network Reconfiguration Using Deep Reinforcement LearningabstractModern networks require frequent configuration updates due to dynamic events, like network expansion and evolving traffic patterns. Existing network reconfiguration tools are effective in certain scenarios, but their practical deployment still has several limitations: (i) They are restricted to specific network topologies, protocols and specifications; (ii) They can cause transient violations; (iii) Their practical deployment is limited by huge computational overheads and the specialized hardware support. To address these limitations, this paper presents a Multi-Objective Optimizing Network Reconfiguration (MONR) framework, which comprises a translator and an optimizer, to automatically generate reconfiguration command sequences. The translator represents various input types into a unified graph format based on Datalog-like facts, which are regardless of the input format. Therefore, MONR supports diverse routing protocols and network specifications. The optimizer employs deep reinforcement learning techniques to simultaneously maximize the specification satisfaction of intermediate configurations and minimize both traffic shifts and the number of command updates to reduce computational overheads. It models the reconfiguration task as a multi-objective Markov Decision Process (MOMDP) and introduces a Dueling Prioritized Experience Replay Double Deep Q-Network (DPER-DDQN) algorithm to balance multiple objectives. We compare MONR with Snowcap, AED and ConfigReco. The evaluation demonstrates that MONR is 2x, 9x, and 56x faster than Snowcap, AED and ConfigReco. Furthermore, MONR maintains 100% specification consistency while reducing the traffic shifts (< 0.1) and the number of update commands (0.8 of Snowcap's). Yuqi Dai, Hua Zhang 0002, Jingyu Wang 0001, Jianxin Liao |
ICNP | 4 |
| 2024 | Following the Compass: LLM-Empowered Intent Translation with Manual GuidanceabstractIntent-Based Networking (IBN) represents a novel paradigm of network automation and intelligence that has gradually been applied to network management. While the emergence of Large Language Models (LLMs) has improved the current state of IBN, hardware heterogeneity and high network dynamics remain significant challenges. Hardware heterogeneity requires that IBN effectively manage a diverse range of devices. The high network dynamics demands that IBN align service needs with rapidly changing network resources. We propose LIT, a framework of LLM-empowered Intent Translation with manual guidance. Given the outstanding language understanding and generation capabilities of LLM, LIT utilizes it in intent translation task. To further address two prevalent problems encountered in IBN, we introduce manual guidance and Mixture of Experts (MoE). Under the guidance of the manual, LLM improves its ability to generate high-quality policies that comply with syntax. After introducing MoE, it makes fine-grained adjustments to the parameters of policies based on network status and service requirements. The experimental outcomes demonstrate that LIT considerably alleviates numerous current challenges confronted by IBN and excels in intent translation, attaining an F1 score that is$\mathbf{5 6. 7 \%}$higher than the baseline model. Lingqi Guo, Jingyu Wang 0001, Caijun Yan, Haifeng Sun 0001, Zirui Zhuang, Qi Qi 0001, Haibao Ren, Jianxin Liao |
ICNP | 10 |
| 2024 | Safeguarding Sustainable Cities: Unsupervised Video Anomaly Detection through Diffusion-based Latent Pattern Learning
Menghao Zhang 0004, Jingyu Wang 0001, Qi Qi 0001, Pengfei Ren 0001, Haifeng Sun 0001, Zirui Zhuang, Lei Zhang 0094, Jianxin Liao |
IJCAI | 8 |
| 2024 | Work Together to Keep Fresh: Hierarchical Learning for UAVs-assisted Data Time-Sensitive IoTabstractIn the context of disaster warning and monitoring within the Internet of Things (IoT), the utilization of unmanned aerial vehicles (UAVs) as relays to gather time-sensitive data from disaster monitoring sensors and transmit it to the base station (BS) has emerged as a highly promising application. In UAV-assisted Data Time-Sensitive IoT (DTIoT), the Age of Information (AoI) serves as a critical performance metric that quantifies the timeliness of data collection, specifically referring to the duration it takes for data to travel from the sensor to the BS. Controlling the flight trajectories of multiple UAVs to minimize AoI is a challenge under energy constraints. Existing work typically uses deep reinforcement learning (DRL) algorithms to address UAV trajectory control problems. However, the task of controlling multi-agent continuous trajectories in complex DTIoT network states is hindered by sparse rewards, posing challenges in training deep neural network-based control policies using standard DRL methods. In this paper, we propose a network-oriented hierarchical reinforcement learning (NO-HRL) algorithm to control the UAVs’ flight trajectory in DTIoT networks for minimizing the AoI. We design the control policy based on a two-layer hierarchical DRL, where the upper layer selects the target and the lower layer executes it. We further propose a decoupled sequential training scheme for effectively training the mutually coupled two-layer DRL network of NO-HRL. The experiment results show that our algorithm outperforms other baselines in AoI optimization for DTIoT. Dezhi Chen, Qianlong Fu, Qi Qi 0001, Jingyu Wang 0001, Jianxin Liao |
IJCNN | 6 |
| 2024 | Fast Policy Convergence for Traffic Engineering with Proactive Distributed Message-PassingabstractNowadays, the rise of various network applications makes network traffic become increasingly complex, which brings more stringent requirements to traffic engineering (TE). Although the state-of-the-art TE approaches based on deep reinforcement learning (DRL) or traditional methods can generate optimal solutions for fixed traffic matrices, they cannot converge fast enough to provide real-time optimization in real networks either because of excessive computation times or high communication overheads. Moreover, due to the dynamically changing traffic load on the network, it is also challenging to achieve optimization of maximum link utilization (MLU) and end-to-end delay at the same time since these two optimization objectives may be conflicting, especially when the network is under a low traffic load, which makes the modeling very difficult. To meet these challenges, we present RT-TE, a TE system based on DRL and distributed message-passing between intelligent agents that can achieve real-time optimization for both MLU and end-to-end delay. To reduce the communication time due to link propagation delay during the optimization process, we design a proactive message-passing mechanism that allows agents to use partial messages to compute the routing policy while maintaining the optimization performance. Additionally, to achieve the tradeoff between the two optimization objectives, we model the propagation delay into the DRL model and design a multi-objective training framework with parameter transfer for training. Based on theoretical modeling, we can find the best tradeoff between the two objectives. Moreover, to improve the model's generalization for various traffic flows, we use a GNN model to generate the rewards of the DRL model, which greatly speeds up the training phase and allows us to feed massive amounts of data into the model. Through evaluations of real-world network topologies, our approach shows a 10%-20% improvement in optimizing MLU under short traffic-changing intervals and yields a 9%-13% improvement in optimizing end-to-end delay compared to state-of-the-art approaches. Zirui Zhuang, Jingyu Wang 0001, Qi Qi 0001, Haifeng Sun 0001, Jianxin Liao |
IPDPS | 6 |
| 2024 | Beyond Throughput-Optimal: Second-Order Smooth Backpressure Algorithm for Reducing Jitter and DelayabstractIn the imminent era of 6G, Quality of Service (QoS) emerges as a pivotal concern in wireless communications. The prescribed transmission rates and vast access demands mandated by 6G standards impose heightened requirements on network throughput and delay. However, the highly dynamic and often bursty nature of application demands presents challenges for routing and congestion control. The backpressure-based joint rate and routing control algorithm adaptively adjusts network traffic to achieve optimal throughput. However, varying traffic conditions hinder the algorithm’s convergence to ideal states. Additionally, relying solely on first-order backlog differences for forwarding can lead to poor convergence and high delays. In this study, we propose a Second-Order Smooth Backpressure (SoSBP) algorithm, leveraging second-order backlog metrics and dual-level queue mapping, to address throughput, delay, and jitter issues in dynamic network environments. We validate the efficacy of this novel backlog metric using Lyapunov optimization techniques. Simulation results demonstrate that our approach significantly reduces end-to-end delay and data jitter while preserving throughput and eliminating routing loops. Yuexi Yin, Zirui Zhuang, Jingyu Wang 0001, Qi Qi 0001, Haifeng Sun 0001, Xiaoyuan Fu, Jianxin Liao |
IWQoS | 7 |
| 2024 | Video Anomaly Detection via Progressive Learning of Multiple Proxy TasksabstractLearning multiple proxy tasks is a popular training strategy in semi-supervised video anomaly detection. However, the traditional method of learning multiple proxy tasks simultaneously is prone to suboptimal solutions, and simply executing multiple proxy tasks sequentially cannot ensure continuous performance improvement. In this paper, we thoroughly investigate the impact of task composition and training order on performance enhancement. We find that ensuring continuous performance improvement in multi-task learning requires different but continuous optimization objectives in different training phases. To this end, a training strategy based on progressive learning is proposed to enhance the multi-task learning in VAD. The learning objectives of the model in previous phases contribute to the training in subsequent phases. Specifically, we decompose video anomaly detection into three phases: perception, comprehension, and inference, continuously refining the learning objectives to enhance model performance. In the three phases, we perform the visual task, the semantic task and the open-set task in turn to train the model. The model learns different levels of features and focuses on different types of anomalies in different phases. Extensive experiments demonstrate the effectiveness of our method, highlighting that the benefits derived from the progressive learning transcend specific proxy tasks. Menghao Zhang 0004, Jingyu Wang 0001, Qi Qi 0001, Pengfei Ren 0001, Haifeng Sun 0001, Zirui Zhuang, Huazheng Wang, Lei Zhang 0094, Jianxin Liao |
ACM Multimedia | 9 |
| 2024 | QUIC-Enabled Framework for Alleviating Transient Congestion in Time-Critical IoTabstractThe real-time control capability of IoT devices is contingent upon the transmission of packets. However, due to the influence of multiple devices accessing the network, the bandwidth available to IoT devices from access points may decline significantly, which causes a surge in the queuing latency and interrupts the transmission. This phenomenon is referred to as transient congestion. To achieve stable high-quality network service, this paper designs RushWay, a QUIC-enabled framework for alleviating transient congestion. RushWay employs stream multiplexing to compress the original packet into the stream frame, thereby reducing the bandwidth required for transmission. Furthermore, RushWay employs an adaptive decision-making algorithm to assess uplink queue conditions and packet latency requirements, thereby alleviating transient congestion. The simulation results demonstrate that RushWay can improve key performance by 13% to 94%. Bo He 0003, Jingyu Wang 0001, Qi Qi 0001, Haifeng Sun 0001, Zirui Zhuang, Jianxin Liao |
MobiCom | 7 |
| 2024 | MDR: Model-Specific Demonstration Retrieval at Inference Time for In-Context LearningabstractHuazheng Wang, Jinming Wu, Haifeng Sun, Zixuan Xia, Daixuan Cheng, Jingyu Wang, Qi Qi, Jianxin Liao. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Huazheng Wang, Haifeng Sun 0001, Zixuan Xia, Daixuan Cheng, Jingyu Wang 0001, Qi Qi 0001, Jianxin Liao |
NAACL-HLT | 8 |
| 2024 | FM-Delta: Lossless Compression for Storing Massive Fine-tuned Foundation ModelsabstractPre-trained foundation models, particularly large language models, have achieved remarkable success and led to massive fine-tuned variants. These models are commonly fine-tuned locally and then uploaded by users to cloud platforms such as HuggingFace for secure storage. However, the huge model number and their billion-level parameters impose heavy storage overhead for cloud with limited resources. Our empirical and theoretical analysis reveals that most fine-tuned models in cloud have a small difference (delta) from their pre-trained models. To this end, we propose a novel lossless compression scheme FM-Delta specifically for storing massive fine-tuned models in cloud. FM-Delta maps fine-tuned and pre-trained model parameters into integers with the same bits, and entropy codes their integer delta. In this way, cloud only needs to store one uncompressed pre-trained model and other compressed fine-tuned models.
Extensive experiments have demonstrated that FM-Delta efficiently reduces cloud storage consumption for massive fine-tuned models by an average of around 50% with only negligible additional time in most end-to-end cases. For example, on up to 10 fine-tuned models in the GPT-NeoX-20B family, FM-Delta reduces the original storage requirement from 423GB to 205GB, significantly saving cloud storage costs. Wanyi Ning, Jingyu Wang 0001, Qi Qi 0001, Mengde Zhu, Haifeng Sun 0001, Daixuan Cheng, Jianxin Liao |
NeurIPS | 7 |
| 2024 | Rethinking the Power of Timestamps for Robust Time Series Forecasting: A Global-Local Fusion PerspectiveabstractTime series forecasting has played a pivotal role across various industries, including finance, transportation, energy, healthcare, and climate. Due to the abundant seasonal information they contain, timestamps possess the potential to offer robust global guidance for forecasting techniques. However, existing works primarily focus on local observations, with timestamps being treated merely as an optional supplement that remains underutilized. When data gathered from the real world is polluted, the absence of global information will damage the robust prediction capability of these algorithms. To address these problems, we propose a novel framework named GLAFF. Within this framework, the timestamps are modeled individually to capture the global dependencies. Working as a plugin, GLAFF adaptively adjusts the combined weights for global and local information, enabling seamless collaboration with any time series forecasting backbone. Extensive experiments conducted on nine real-world datasets demonstrate that GLAFF significantly enhances the average performance of widely used mainstream forecasting models by 12.5\%, surpassing the previous state-of-the-art method by 5.5\%. Chengsen Wang, Qi Qi 0001, Jingyu Wang 0001, Haifeng Sun 0001, Zirui Zhuang, Jianxin Liao |
NeurIPS | 7 |
| 2024 | EPVerifier: Accelerating Update Storms Verification with Edge-Predicate
Chenyang Zhao 0005, Yuebin Guo, Jingyu Wang 0001, Qi Qi 0001, Zirui Zhuang, Haifeng Sun 0001, Lingqi Guo, Yuming Xie, Jianxin Liao |
NSDI | 9 |
| 2024 | INCS: Intent-driven network-wide configuration synthesis based on deep reinforcement learning
Yuqi Dai, Hua Zhang 0002, Jingyu Wang 0001, Jianxin Liao |
Comput. Networks | 4 |
| 2024 | HierNet: A Hierarchical Resource Allocation Method for Vehicle Platooning NetworksabstractVehicle platooning is a promising traffic model in intelligent transportation systems (ITSs), which can effectively improve resource utilization and reduce traffic congestion. The resource allocation for vehicle-to-everything (V2X) communications that consist of intraplatoon communications and interplatoon communications is crucial for safe operation of multiple vehicular platoons. Considering dynamic coordination pattern of vehicular platoons and layered architecture of vehicle platooning networks, a hierarchical resource decision-making framework is proposed in this article. In the proposed framework, the resource decision-making process is divided into two levels. The high level that generates and distributes coordination meta policy is deployed on base station (BS), and the low level that generates ego resource decisions is deployed in each platoon. To deal with optimization of resource allocation for multiplatoon V2X communications, a hierarchical reinforcement learning method (HierNet) is designed based on the proposed hierarchical decision-making framework. In HierNet, meta policy of the high level can be preserved and needs to be updated only when cooperative conditions of multiple platoons undergo distinct changes. Simulation experiments have demonstrated that our proposed method not only optimizes resource efficiency but also reduces the communication costs for resource decision making of vehicle platooning networks. Xiaoyuan Fu, Quan Yuan 0004, Guiyang Luo, Nan Cheng 0001, Jianxin Liao |
IEEE Internet Things J. | 7 |
| 2024 | TacNet: A Tactic-Interactive Resource Allocation Method for Vehicular NetworksabstractTo support safety driving and various on-board services, efficient resource allocation is crucial for the promising implement of vehicle platooning in intelligent transportation systems (ITSs). The resource allocation of vehicle-to-everything (V2X) communications for vehicular platoons is studied in this article. First, a multiobjective function is formulated to jointly optimize sub-band and power allocation to satisfy Quality-of- Service (QoS) in vehicular networks. With the advantage of dealing with complex decision-making problems in multiagent systems, distributed multiagent deep reinforcement learning (MADRL) stands out for resource allocation of vehicular networks. However, it faces the challenge of cooperation aging when every agent is only learning from information of others to form a cooperation model in the training process. Considering the random and dynamic combination of vehicles in vehicle platooning, a tactic-interactive MADRL method named as TacNet is then proposed to improve the cooperation efficiency of multiple agents. In TacNet, the tactics of other agents will be encoded and transmitted through interactive communications among agents. In addition, with the development of vehicular edge computing (VEC), digital twin (DT) networks are constructed to assist offloading computation-intensive resource allocation tasks in vehicles to the edge. The superiority of the proposed method is verified through extensive simulation results, which refers to convergence and performance of satisfying diversified QoS requirements compared with state-of-the-art MADRL methods. Xiaoyuan Fu, Quan Yuan 0004, Zirui Zhuang, Jianxin Liao, Dongmei Zhao |
IEEE Internet Things J. | 5 |
| 2024 | SMR: Spatial-Guided Model-Based Regression for 3D Hand Pose and Mesh Reconstructionabstract3D hand reconstruction is an important technique for human-computer interaction. Interactive experience depends on the accuracy, efficiency, and robustness of the algorithm. Therefore, in this paper, we first propose a balanced framework called spatial-aware regression (SAR) to achieve precise and fast reconstruction. SAR can bridge convolutional networks and graph-structure networks more effectively than existing frameworks to fully exploit extracted spatial information using a novel spatial-aware initial graph building module. In addition, SAR uses adaptive-GCN to make keypoints interact efficiently and effectively; and regresses 2.5D belief maps to characterize uncertainty. SAR is highly flexible because it can predict an arbitrary number of keypoints and apply pose-guided refinement for coarse to fine regression. To produce more rational results for challenging cases and mitigate 3D label reliance, we also propose a more robust model-based framework called spatial-guided model-based regression (SMR) that is based on SAR. There are two critical designs of SMR: 1) it uses SAR to enhance the features with pose information to help the regression of hand model parameters; and 2) it regresses parameters in a spatially aware manner that is similar to SAR. Experiments demonstrate that the proposed frameworks surpass existing fully-supervised approaches on the FreiHAND, HO-3D, RHD, and STB datasets. Also, the performances of the proposed frameworks under weakly/self-supervised settings outperform other competitors. Meanwhile, the proposed frameworks are accurate and efficient. Haifeng Sun 0001, Xiaozheng Zheng, Pengfei Ren 0001, Jingyu Wang 0001, Qi Qi 0001, Jianxin Liao |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Transformer-Based Reinforcement Learning for Scalable Multi-UAV Area CoverageabstractCompared with terrestrial networks, unmanned aerial vehicles (UAVs) have the characteristics of flexible deployment and strong adaptability, which are an important supplement to intelligent transportation systems (ITS). In this paper, we focus on the multi-UAV network area coverage problem (ACP) which require intelligent UAVs long-term trajectory decisions in the complex and scalable network environment. Multi-agent deep reinforcement learning (DRL) has recently emerged as an effective tool for solving long-term decisions problems. However, since the input dimension of multi-layer perceptron (MLP)-based deep neural network (DNN) is fixed, it is difficult for standard DNN to adapt to a variable number of UAVs and network users. Therefore, we combine Transformer with DRL to meet the scalability of the network and propose a Transformer-based deep multi-agent reinforcement learning (T-MARL) algorithm. Transformer can adapt to variable input dimensions and extract important information from complex network states by attention module. In our research, we find that random initialization of Transformer may cause DRL training failure, so we propose a baseline-assisted pre-training scheme. This scheme can quickly provide an initial policy model for UAVs based on imitation learning, and use the temporal-difference(1) algorithm to initialize policy evaluation network. Finally, based on parameter sharing, T-MARL is applicable to any standard DRL algorithm and supports expansion on networks of different sizes. Experimental results show that T-MARL can make UAVs have cooperative behaviors and perform outstandingly on ACP. Dezhi Chen, Qi Qi 0001, Qianlong Fu, Jingyu Wang 0001, Jianxin Liao, Zhu Han 0001 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2024 | ShuttleBus: Dense Packet Assembling With QUIC Stream Multiplexing for Massive IoTabstractIn this paper, we investigate dense short packet forwarding for clustering-based massive Internet-of-Things (mIoT). The objective is to support the data forwarding with minimal communication overhead while satisfying the differentiated latency constraints from the transport layer perspective. To this end, we propose a dense packet assembling scheme, named ShuttleBus, for forwarding devices in mIoT to achieve effective data merging. The assembling scheme is designed based on the stream multiplexing mechanism of the Quick UDP Internet Connection (QUIC) protocol. With ShuttleBus, the payload data sent from IoT devices are extracted as independent frames belonging to different data streams. The ShuttleBus can bundle data frames from multiple streams into a single packet while ensuring data integrity of these streams. Furthermore, we develop a resilient packing mechanism in packet assembling to merge data received from IoT devices within a cluster. In addition, a latency-oriented scheduling mechanism for backlogged QUIC data is established to guarantee satisfactory delivery of diverse transmission tasks. To accommodate the dynamic network environment, we tailor a learning-based algorithm to determine the optimal packet assembling time adaptively. We evaluate the performance of ShuttleBus under various network load conditions. Both analytical and experimental results demonstrate that the proposed scheme significantly reduces communication overhead and enhances data delivery performance under stringent latency constraints. Bo He 0003, Jingyu Wang 0001, Qi Qi 0001, Qiang Ye 0002, Qihao Li, Jianxin Liao, Xuemin Shen |
IEEE Trans. Mob. Comput. | 6 |
| 2024 | Slice Sandwich: Jagged Slicing Multi-Tier Dynamic Resources for Diversified V2X ServicesabstractWith the advancement of intelligent transportation systems, a series of diversified V2X applications come into being, which have different key performance indicators (KPIs) and transmission features. Moreover, multi-tier computing as a new system-level architecture distributes computing and communication capabilities anywhere between the cloud and the end-user. Unfortunately, the existing network paradigm for V2X services adopts a one-shot allocation of resources ignoring the inherent differences of V2X service. To cope with these problems, three types of refined network slices for V2X services are first proposed to simultaneously support heterogeneous service characteristics without excessively splitting resources. Considering the spatiotemporal correlation between service traffic and physical resources, a jagged slicing in multi-tier dynamic resources, which forms a “slice sandwich” brightly, is realized by a dual timescale intelligent resource management scheme. The inter-slice resource configuration is based on neural bandits with upper confidence bounds at each large-time period, while the exclusive resources are managed elastically by deep Q-learning in terms of the real-time changing network state in the small slot. We developed a simulation environment by Simulation of Urban Mobility (SUMO) including real-world road conditions and traffic models. The experiment results demonstrate that the proposed scheme can effectively guarantee KPIs of V2X services and improve the system revenue compared with benchmark algorithms. Yu Liu 0016, Zirui Zhuang, Qi Qi 0001, Jingyu Wang 0001, Dezhi Chen, Lu Lu 0015, Jianxin Liao, Zhu Han 0001 |
IEEE Trans. Mob. Comput. | 8 |
| 2024 | PICO: Pipeline Inference Framework for Versatile CNNs on Diverse Mobile DevicesabstractDistributing the inference of convolutional neural network (CNN) to multiple mobile devices has been studied in recent years to achieve real-time inference without losing accuracy. However, how to map CNN to devices remains a challenge. On the one hand, scheduling the workload of state-of-the-art CNNs with multiple devices is NP-Hard because the structures of CNNs are directed acyclic graphs (DAG) rather than simple chains. On the other hand, distributing the inference workload suffers from expensive communication and unbalanced computation due to the wireless environment and heterogeneous devices. This paper presents PICO, a pipeline cooperation framework to accelerate the inference of versatile CNNs on diverse mobile devices. At its core, PICO features: (1) a generic graph partition algorithm that considers the characteristics of any given CNN and orchestrates it into a list of model pieces with suitable granularity, and (2) a many-to-many mapping algorithm that produces the best pipeline configuration for heterogeneous devices. In our experiment with$2 \sim 8$Raspberry-Pi devices, the throughput can be improved by$1.8 \sim 6.8 \times$under different CPU frequencies. Qi Qi 0001, Jingyu Wang 0001, Haifeng Sun 0001, Jianxin Liao, Song Guo 0001 |
IEEE Trans. Mob. Comput. | 6 |
| 2024 | Dynamic and Static Representation Learning Network for RecommendationabstractExisting review-based recommendation methods learn a latent representation of user and item from user-generated reviews by a static strategy, which are unable to capture the dynamic evolution of users' interests and the dynamic attraction of items. Here, we propose a dynamic and static representation learning network (DSRLN) to improve the rating prediction accuracy by exploring fine-grained representations of users and items. Specifically, we built DSRLN with a dynamic representation extractor to model the dynamic evolution of users' interests by exploring the inner relations of an interaction sequence, and with a static representation extractor to model the users' intrinsic preferences by learning the semantic coherence and feature strength information from reviews. To identify the different influences of dynamic and static features for different users, a personalized adaptive fusion module was designed using a weighted attention mechanism. Extensive experiments on five real-world datasets from Amazon demonstrated the superiority of the proposed model, and the additional ablation studies verified the effectiveness of the components designed in the DSRLN model. Tongcun Liu, Siyuan Lou, Jianxin Liao, Hailin Feng |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Multimodal Multitask Control Plane Verification FrameworkabstractModern networks are susceptible to configuration errors, such as misconfigurations and policy conflicts due to the complex interactions of diverse devices through various protocols. Control plane verification offers an effective solution to prevent these errors. However, existing tools face several challenges: (i) prolonged verification times, (ii) the verification of only specific policies, and (iii) poor robustness against node and link failures. To address these issues, we propose a control plane verification framework based on a multimodal multitask learning model. This framework enables simultaneous verification of multiple policies directly from various network configuration files. The learning model utilizes modality fusion techniques to capture both topology-related and traffic-related network features. It is trained on datasets augmented with the failure model to enhance robustness against failures. We compare our framework with three state-of-the-art verification tools: Minesweeper, Hoyan, and Tiramisu. Our evaluation shows that our framework is 2600 times faster than Minesweeper, twice as fast as Hoyan, and 19 times faster than Tiramisu, while maintaining 100% verification accuracy. Furthermore, our framework excels in verifying traffic-related network policies and remains effective even under node and link failures. Yuqi Dai, Hua Zhang 0002, Jingyu Wang 0001, Jianxin Liao |
IEEE Trans. Netw. Serv. Manag. | 4 |
| 2024 | AdapINT: A Flexible and Adaptive In-Band Network Telemetry System Based on Deep Reinforcement LearningabstractIn-band Network Telemetry (INT) has emerged as a promising network measurement technology. However, existing network telemetry systems lack the flexibility to meet diverse telemetry requirements and are also difficult to adapt to dynamic network environments. In this paper, we propose AdapINT, a versatile and adaptive in-band network telemetry framework assisted by dual-timescale probes, including long-period auxiliary probes (APs) and short-period dynamic probes (DPs). Technically, the APs collect basic network status information, which is used for the path planning of DPs. To achieve full network coverage, we propose an auxiliary probes path deployment (APPD) algorithm based on the Depth-First-Search (DFS). The DPs collect specific network information for telemetry tasks. To ensure that the DPs can meet diverse telemetry requirements and adapt to dynamic network environments, we apply the deep reinforcement learning (DRL) technique and transfer learning method to design the dynamic probes path deployment (DPPD) algorithm. The evaluation results show that AdapINT can flexibly customize the telemetry system to accommodate diverse requirements and network environments. In latency-aware networks, AdapINT effectively reduces telemetry latency, while in overhead-aware networks, it significantly lowers the control overheads. Hua Zhang 0002, Yibo Pi, Zijian Cao 0005, Jingyu Wang 0001, Jianxin Liao |
IEEE Trans. Netw. Serv. Manag. | 6 |
| 2024 | Dynamic Network Slice for Bursty Edge TrafficabstractEdge network slicing promises better utilization of network resources by dynamically allocating resources on demand. However, addressing the imbalance between slice resources and user demands becomes challenging when complex user behaviors lead to bursty traffic within the edge network. Hence, we propose a comprehensive dynamic slice strategy with two coupled sub-strategies (i) bursty-sensitive slice resource coordination and (ii) proactive demand resource matching to find an optimal balance. For obtaining stable strategies, the edge network with bursty traffic is formulated as a bi-level Lyapunov optimization problem. Then we propose a resource allocation and request redirection (RA-RR) algorithm with polynomial complexity by introducing deep reinforcement learning to guarantee real-time. Specifically, two agents are trained to solve two sub-strategies, and the Lyapunov drift-plus-penalty function is used as the reward to keep queues stable. RA-RR is responsive to fluctuations in demand and realizes an efficient interaction of coupled decision-making. Moreover, a training method based on alternating optimization is designed to ensure convergence of the RA-RR algorithm. Experiments demonstrate that the proposal can maximize network revenue while ensuring the stability of slice services when edge traffic bursts, and has an average improvement of 20.4% compared with comparisons. Rongxin Han, Jingyu Wang 0001, Qi Qi 0001, Dezhi Chen, Zirui Zhuang, Haifeng Sun 0001, Xiaoyuan Fu, Jianxin Liao, Song Guo 0001 |
IEEE/ACM Trans. Netw. | 8 |
| 2024 | Fast and Scalable ACL Policy Solving Under Complex Constraints With Graph Neural NetworksabstractNetwork operators often need to modify Access Control List (ACL) policies to align with to network upgrades. An essential part of the ACL update task is reachability satisfaction. Previous studies formalize reachability requirements as a set of constraints and then use Boolean Satisfiability (SAT) or Satisfiability Modulo Theories (SMT) solvers to search for solutions. However, as today’s networks grow in size and complexity, the constraints derived from the requirements become increasingly complex, leading to an unacceptable time cost to obtain a correct policy. The sluggish updating of ACL policies can affect the properties of a network, such as connectivity and security. This paper presents a novel approach for fast and scalable ACL policy synthesis under complex constraints. We utilize Graph Neural Networks (GNNs) to learn the relations between nodes and reason the solution that satisfies the update requirements. We further integrate global position encoding into the GNN architecture, which allows for better differentiation of nodes in ACL update tasks. Additionally, an enhanced stochastic local search solver is introduced to address incorrect predictions made by the GNN. Experiments on real-world topologies show that GNN saves up$278\times $time costs compared to advanced SAT/SMT solvers on a 125-node network, and this advantage expands with the network size. Furthermore, our model extrapolates well when faced with different requirements and topologies, demonstrating its ability to handle frequent network upgrades. Haifeng Sun 0001, Xingjian Liao, Jingyu Wang 0001, Qi Qi 0001, Zirui Zhuang, Jianxin Liao, Dapeng Oliver Wu |
IEEE/ACM Trans. Netw. | 6 |
| 2024 | Diner: Interpretable Anomaly Detection for Seasonal Time Series in Web ServicesabstractMonitoring and anomaly detection of key performance indicators (KPIs) are crucial for large Internet companies to maintain the reliability of their Web services. Influenced by human behavior and schedules, the KPIs of Web services typically exhibit seasonal characteristics. These characteristics may be complex as different KPIs exhibit differences in trend, multiple periods, and noise behaviors. However, existing anomaly detection methods typically only model one fixed pattern of seasonal KPIs, which may lead to performance degradation when dealing with diverse seasonal KPIs. In this work, we propose a novel anomaly detection model for seasonal KPIs,Diner, which incorporates multiple interpretable components. It is able to capture the additive and multiplicative trends, multiple periods, and seasonal noise in intricate seasonal KPIs, making it easily adaptable to different types of seasonal KPIs. Additionally, we present a set of evaluation criteria for generic time series anomaly detection tasks, which prove more effective in handling ambiguous manual labels and various anomaly events. Experiments are conducted on three real-world datasets, and the performanceDinersurpassed both the statistical baseline and the state-of-the-art deep learning baselines. Yuhan Jing, Jingyu Wang 0001, Ji Qi 0005, Qi Qi 0001, Bo He 0003, Zirui Zhuang, Naixing Wu, Jianxin Liao |
IEEE Trans. Serv. Comput. | 8 |
| 2024 | One Teacher is Enough: A Server-Clueless Federated Learning With Knowledge DistillationabstractMachine learning-based services offer intelligent solutions with powerful models. To enhance model robustness, Federated Learning (FL) emerges as a promising collaborative learning paradigm, which iteratively trains a global model through parameter exchange among multiple clients based on their local data. Generally, the local data are heterogeneous, which slows down convergence. Knowledge distillation is an effective technique against data heterogeneity while existing works distill the ensemble knowledge from local models, ignoring the natural global knowledge from the aggregated model. This places limitations on their algorithms, such as the need for proxy data or the necessary exposure of local models to the server, which is prohibited in most privacy-preserving FL with a clueless server. In this work, we propose FedDGT, a novel knowledge distillation method for industrial server-clueless FL. FedDGT regards the aggregated model as the only one teacher to impart its global knowledge into a generator and then regularizes the drifted local models through the generator, overcoming previous limitations and providing better privacy and scalability support. Extensive experiments demonstrate that FedDGT can achieve highly-competitive model performance while greatly reducing the communication rounds in a server-clueless scenario. Wanyi Ning, Qi Qi 0001, Jingyu Wang 0001, Mengde Zhu, Shaolong Li, Jianxin Liao |
IEEE Trans. Serv. Comput. | 7 |
| 2024 | PsyQoE: Improving Quality-of-Experience Assessment With Psychological Effects in Video StreamingabstractAs media streaming grows, service providers have to focus more on users’ quality of experience (QoE). Among the metrics that influence QoE, network-oriented and media-oriented factors have been the primary concern of most existing work. Apart from the two factors, however, another crucial aspect of QoE isthe user. Though the analysis and modeling of users is a turning point in the transition from QoS to QoE, the impact of users’ memory and cognitive-psychological effects have not been fully explored. In this work, we analyze, validate and quantify the relationship between multiple cognitive effects and the QoE of users in media streaming. We propose PsyQoE, a new QoE assessment framework that uniquely focuses on user perceptions during media sessions. Besides providing real-time QoE assessments through leveraging machine learning, PsyQoE also predicts the overall QoE by taking into account user biases and long-term effects. Additionally, we have designed a further modified version to cater to the needs of hardware and energy-restricted scenarios. When compared with other existing methods, our approach demonstrates enhanced performance by boosting accuracy by 7% to 16%. It also provides better interpretability and broader applicability. Daoxu Sheng, Qi Qi 0001, Jingyu Wang 0001, Lianyuan Li, Jianxin Liao |
IEEE Trans. Serv. Comput. | 6 |
| 2024 | Cognition Guided Video Anomaly Detection Framework for Surveillance ServicesabstractThe aim of surveillance services is to detect anomalous events that occur in given surveillance videos. Most existing video anomaly detection methods rely on minimizing reconstruction or prediction errors due to the lack of abnormal data, which results in poor generalization and overfitting. In fact, cognitions for anomalies in surveillance videos mainly relies on crucial relationships, including ones between objects and ones between objects and scenes. Focusing on this property of anomaly detection, aCognitionGuidedVideoAnomalyDetection framework based on prior knowledge is proposed, calledCG-VAD. CG-VAD introduces both explicit and implicit prior knowledge into the frame prediction network to let the model exploit crucial relationships. Explicit knowledge containing crucial relationships related to anomaly is introduced into the anomaly detection model through a proposed embedding network based on multi-layer Graph Convolutional Networks. Implicit knowledge in the form of learnable parameters enhances the ability of the model to learn crucial relationships through prompt tuning. By integrating prior knowledge to focus the model on the relationships associated with the anomaly, we find that CG-VAD is not only quick to adapt to new real-world scenarios, but it is also able to recognize the type of anomaly. We have conducted extensive experiments on four benchmark datasets and the results indicate that the proposed method outperforms previous methods. Specifically, CG-VAD achieves an AUROC score of 87.2$\%$on the ShanghaiTech dataset. Code is available athttps://github.com/zmh0124/CG-VAD. Menghao Zhang 0004, Jingyu Wang 0001, Qi Qi 0001, Zirui Zhuang, Haifeng Sun 0001, Jianxin Liao |
IEEE Trans. Serv. Comput. | 6 |
| 2023 | Scene-Level Sketch-Based Image Retrieval with Minimal Pairwise SupervisionabstractThe sketch-based image retrieval (SBIR) task has long been researched at the instance level, where both query sketches and candidate images are assumed to contain only one dominant object. This strong assumption constrains its application, especially with the increasingly popular intelligent terminals and human-computer interaction technology. In this work, a more general scene-level SBIR task is explored, where sketches and images can both contain multiple object instances. The new general task is extremely challenging due to several factors: (i) scene-level SBIR inherently shares sketch-specific difficulties with instance-level SBIR (e.g., sparsity, abstractness, and diversity), (ii) the cross-modal similarity is measured between two partially aligned domains (i.e., not all objects in images are drawn in scene sketches), and (iii) besides instance-level visual similarity, a more complex multi-dimensional scene-level feature matching problem is imposed (including appearance, semantics, layout, etc.). Addressing these challenges, a novel Conditional Graph Autoencoder model is proposed to deal with scene-level sketch-images retrieval. More importantly, the model can be trained with only pairwise supervision, which distinguishes our study from others in that elaborate instance-level annotations (for example, bounding boxes) are no longer required. Extensive experiments confirm the ability of our model to robustly retrieve multiple related objects at the scene level and exhibit superior performance beyond strong competitors. Ce Ge, Jingyu Wang 0001, Qi Qi 0001, Haifeng Sun 0001, Tong Xu 0002, Jianxin Liao |
AAAI | 6 |
| 2023 | Semi-transductive Learning for Generalized Zero-Shot Sketch-Based Image RetrievalabstractSketch-based image retrieval (SBIR) is an attractive research area where freehand sketches are used as queries to retrieve relevant images. Existing solutions have advanced the task to the challenging zero-shot setting (ZS-SBIR), where the trained models are tested on new classes without seen data. However, they are prone to overfitting under a realistic scenario when the test data includes both seen and unseen classes. In this paper, we study generalized ZS-SBIR (GZS-SBIR) and propose a novel semi-transductive learning paradigm. Transductive learning is performed on the image modality to explore the potential data distribution within unseen classes, and zero-shot learning is performed on the sketch modality sharing the learned knowledge through a semi-heterogeneous architecture. A hybrid metric learning strategy is proposed to establish semantics-aware ranking property and calibrate the joint embedding space. Extensive experiments are conducted on two large-scale benchmarks and four evaluation metrics. The results show that our method is superior over the state-of-the-art competitors in the challenging GZS-SBIR task. Ce Ge, Jingyu Wang 0001, Qi Qi 0001, Haifeng Sun 0001, Tong Xu 0002, Jianxin Liao |
AAAI | 6 |
| 2023 | Two Heads Are Better than One: Image-Point Cloud Network for Depth-Based 3D Hand Pose EstimationabstractDepth images and point clouds are the two most commonly used data representations for depth-based 3D hand pose estimation. Benefiting from the structuring of image data and the inherent inductive biases of the 2D Convolutional Neural Network (CNN), image-based methods are highly efficient and effective. However, treating the depth data as a 2D image inevitably ignores the 3D nature of depth data. Point cloud-based methods can better mine the 3D geometric structure of depth data. However, these methods suffer from the disorder and non-structure of point cloud data, which is computationally inefficient. In this paper, we propose an Image-Point cloud Network (IPNet) for accurate and robust 3D hand pose estimation. IPNet utilizes 2D CNN to extract visual representations in 2D image space and performs iterative correction in 3D point cloud space to exploit the 3D geometry information of depth data. In particular, we propose a sparse anchor-based "aggregation-interaction-propagation'' paradigm to enhance point cloud features and refine the hand pose, which reduces irregular data access. Furthermore, we introduce a 3D hand model to the iterative correction process, which significantly improves the robustness of IPNet to occlusion and depth holes. Experiments show that IPNet outperforms state-of-the-art methods on three challenging hand datasets. Pengfei Ren 0001, Jiachang Hao, Haifeng Sun 0001, Qi Qi 0001, Jingyu Wang 0001, Jianxin Liao |
AAAI | 7 |
| 2023 | Region-Aware Dynamic Filtering Network for 3D Hand Reconstructionabstract3D hand reconstruction from RGB image has attracted a lot of attention due to its crucial role in human-computer interaction. Nevertheless, it is still challenging to perform 3D hand reconstruction under conditions of hand-object interaction due to severe mutual occlusion. Previous methods usually adopt fixed convolution kernel to extract features. We argue that simply sharing the static filter for all regions is impertinent, given that the occlusion degree varies across different regions, resulting in inconsistent visual representations. To address this issue, we proposed Region-aware Dynamic Filtering Network (RDFNet), which dynamically generates convolution kernels based on the features of different regions, thereby adaptively extracting region-related information. Furthermore, we introduce a dynamic receptive field selection mechanism to determine the most appropriate scale for the convolution kernel. For the severely occluded regions, larger receptive field is needed to capture semantic-related features, while the visible regions are mainly concerned with their own local pattern to accumulate spatial-related features and avoid the interference of irrelevant information. Our proposed RDFNet outperforms state-of-the-art methods by a large margin on several challenging hand-object interaction datasets. Pengfei Ren 0001, Jingyu Wang 0001, Haifeng Sun 0001, Qi Qi 0001, Jing Wang 0039, Jianxin Liao |
ECAI | 7 |
| 2023 | Reasoning Guided by a Manual: Context-Aware Image Captioning with Novel ObjectsabstractNovel object captioning task aims at describing objects that are absent from training data. Due to the scarcity of novel objects, it’s challenging to find a way to utilize external data to improve model’s reasoning ability. While previously designed methods all follow a deep learning approach, we boost novel object captioning by incorporating reasoning with traditional deep learning framework. We design a manual from dictionaries that provides our model with sufficient and accurate external information on novel objects. We propose Manual-guided Context-aware Novel Object Captioning model (MC-NOC) that utilizes image and caption context to generate novel object captions. It contains a Manual-Guided Novel Object Reasoning module to reason about novel objects based on other objects of the given image and a Caption Reconstruction module to incorporate novel objects into generated captions according to caption context. We validate MC-NOC with state-of-the-art performance on the challenging Held-out COCO and Nocaps dataset, leading their leaderboard. In particular, we improved the CIDER metric by 6.4 points on the held-out coco dataset. Comprehensive experiments demonstrate our model’s reasoning capability and the quality of generated captions. Peiyao Hua, Haifeng Sun 0001, Jiachang Hao, Cong Liu 0046, Jingyu Wang 0001, Qi Qi 0001, Jianxin Liao |
ECAI | 7 |
| 2023 | How Does Diffusion Influence Pretrained Language Models on Out-of-Distribution Data?abstractTransformer-based pretrained language models (PLMs) have achieved great success in modern NLP. An important advantage of PLMs is good out-of-distribution (OOD) robustness. Recently, diffusion models have attracted a lot of work to apply diffusion to PLMs. It remains under-explored how diffusion influences PLMs on OOD data. The core of diffusion models is a forward diffusion process which gradually applies Gaussian noise to inputs, and a reverse denoising process which removes noise. The noised input reconstruction is a fundamental ability of diffusion models. We directly analyze OOD robustness by measuring the reconstruction loss, including testing the abilities to reconstruct OOD data, and to detect OOD samples. Experiments are conducted by analyzing different training parameters and data statistical features on eight datasets. It shows that finetuning PLMs with diffusion degrades the reconstruction ability on OOD data. The comparison also shows that diffusion models can effectively detect OOD samples, achieving state-of-the-art performance in most of the datasets with an absolute accuracy improvement up to 18%. These results indicate that diffusion reduces OOD robustness of PLMs. Huazheng Wang, Daixuan Cheng, Haifeng Sun 0001, Jingyu Wang 0001, Qi Qi 0001, Jianxin Liao, Jing Wang 0039, Cong Liu 0046 |
ECAI | 6 |
| 2023 | Turn on the Right Track: Weakly Supervised Video Moment Retrieval with Self-Improving Query ReconstructionabstractExisting weakly-supervised temporal sentence grounding methods typically regard query reconstruction as the pretext task in place of the absent temporal supervision. However, their approaches suffer from two flaws, i.e. insignificant reconstruction and discrepancy in alignment. Insignificant reconstruction indicates the randomly masked words may not be discriminative enough to distinguish the target event from unrelated events in the video. Discrepancy in alignment indicates the incorrect partial alignment built by query reconstruction task. The flaws undermine the reliability of current reconstruction-based methods. To this end, we propose a novel Self-improving Query ReconstrucTion (SQRT) framework for weakly-supervised temporal sentence grounding. To deal with insignificant reconstruction, we devise a key words mining strategy to determine the important words for language grounding. To attain better moment-query alignment, we introduce inter-sample contrast to tackle the partial alignment built by query reconstruction. The self-improving framework utilizes query reconstruction for language grounding and alleviates the discrepancy in alignment, thus turning on the right track. Experiments on two popular datasets show that SQRT achieves state-of-the-art performance on Charades-STA and comparable performance to the state-of-the-art on ActivityNet Captions. Haifeng Sun 0001, Jiachang Hao, Jing Wang 0039, Qi Qi 0001, Jingyu Wang 0001, Jianxin Liao |
ECAI | 8 |
| 2023 | Decoupled Iterative Refinement Framework for Interacting Hands Reconstruction from a Single RGB ImageabstractReconstructing interacting hands from a single RGB image is a very challenging task. On the one hand, severe mutual occlusion and similar local appearance between two hands confuse the extraction of visual features, resulting in the misalignment of estimated hand meshes and the image. On the other hand, there are complex spatial relationship between interacting hands, which significantly increases the solution space of hand poses and increases the difficulty of network learning. In this paper, we propose a decoupled iterative refinement framework to achieve pixel-alignment hand reconstruction while efficiently modeling the spatial relationship between hands. Specifically, we define two feature spaces with different characteristics, namely 2D visual feature space and 3D joint feature space. First, we obtain joint-wise features from the visual feature map and utilize a graph convolution network and a transformer to perform intra- and inter-hand information interaction in the 3D joint feature space, respectively. Then, we project the joint features with global information back into the 2D visual feature space in an obfuscation-free manner and utilize the 2D convolution for pixel-wise enhancement. By performing multiple alternate enhancements in the two feature spaces, our method can achieve an accurate and robust reconstruction of interacting hands. Our method outperforms all existing two-hand reconstruction methods by a large margin on the InterHand2.6M dataset. Pengfei Ren 0001, Xiaozheng Zheng, Zhou Xue, Haifeng Sun 0001, Qi Qi 0001, Jingyu Wang 0001, Jianxin Liao |
ICCV | 8 |
| 2023 | CONFPILOT: A Pilot for Faster Configuration by Learning from Device ManualsabstractThe command line interface (CLI) is widely used to configure and manage network devices. However, as heterogeneous devices are introduced into the network, the CLI-based method is becoming time-consuming and inefficient because much effort is required to learn proprietary configuration languages of different vendors or consult online documents. In this work, we present CONFPILOT, an assistant system that can accelerate configuration by automatically converting natural language intents into commands. Our solution is based on a retrieval-augmented generation framework that features a unified parser that parses device manuals into a searchable configuration library, a vendor-agnostic retriever that finds the most relevant$k$syntaxes through a two-stage coarse-to-fine process, and a reliable generator that predicts syntactically correct commands via a pointer-generator network and syntax-guided decoding. In a nutshell, CONFPILOT frees engineers from most time-consuming efforts by learning directly from device manuals to generate configuration commands. Our evaluation and user study show, CONFPILOT can speed up the configuration process by 60x compared to manual lookup while maintaining acceptable exact match accuracy. Furthermore, CONFPILOT can quickly adapt to new vendors and devices with little human effort and time cost. Jinyu Zhao, Haifeng Sun 0001, Jingyu Wang 0001, Qi Qi 0001, Zirui Zhuang, Shimin Tao, Jianxin Liao |
ICDCS | 7 |
| 2023 | Solving Distributed ACL Policies Under Complex Constraints with Graph Neural NetworksabstractAccess Control List (ACL) policies often need to be updated due to upgrades in network architecture and services. A critical part of ACL update tasks is reachability satisfaction, which is typically handled using Boolean Satisfiability (SAT) or Satisfiability Modulo Theories (SMT) solvers. However, as modern networks grow in size and complexity, the constraints derived from reachability requirements become increasingly complex, resulting in a considerable time cost to obtain a satisfying policy. The slow update of ACL policies can endanger network connectivity and security. This paper presents a new approach for fast and scalable ACL policy synthesis under complex constraints. We leverage Graph Neural Networks (GNNs) to learn the relations between nodes and reason the solution that satisfies the update requirements. In addition, an enhanced stochastic local search solver is introduced to deal with erroneous predictions of the GNN. Evaluations show that the proposed method guarantees 100% accuracy on real-world topologies. GNN outperforms modern SAT/SMT solvers in speed, saving up to 278x time costs on a 125-node topology. Furthermore, our method extrapolates well when faced with different requirements and topologies. Xingjian Liao, Haifeng Sun 0001, Jingyu Wang 0001, Qi Qi 0001, Zirui Zhuang, Jianxin Liao |
ICNP | 6 |
| 2023 | Not Only Pairwise Relationships: Fine-Grained Relational Modeling for Multivariate Time Series ForecastingabstractRecent graph-based methods achieve significant success in multivariate time series modeling and forecasting due to their ability to handle relationships among time series variables. However, only pairwise relationships are considered in most existing works. They ignore beyond-pairwise relationships and their potential categories in practical scenarios, which leads to incomprehensive relationship learning for multivariate time series forecasting. In this paper, we present ReMo, a Relational Modeling-based method, to promote fine-grained relational learning among multivariate time series data. Firstly, by treating time series variables and complex relationships as nodes and hyperedges, we extract multi-view hypergraphs from data to capture beyond-pairwise relationships. Secondly, a novel hypergraph message passing strategy is designed to characterize both nodes and hyperedges by inferring the potential categories of relationships and further distinguishing their impacts on time series variables. By integrating these two modules into the time series forecasting framework, ReMo effectively improves the performance of multivariate time series forecasting. The experimental results on seven commonly used datasets from different domains demonstrate the superiority of our model. Qi Qi 0001, Jingyu Wang 0001, Haifeng Sun 0001, Zhikang Wu, Zirui Zhuang, Jianxin Liao |
IJCAI | 7 |
| 2023 | Drift doesn't Matter: Dynamic Decomposition with Diffusion Reconstruction for Unstable Multivariate Time Series Anomaly DetectionabstractMany unsupervised methods have recently been proposed for multivariate time series anomaly detection. However, existing works mainly focus on stable data yet often omit the drift generated from non-stationary environments, which may lead to numerous false alarms. We propose **D**ynamic **D**ecomposition with **D**iffusion **R**econstruction (D$^3$R), a novel anomaly detection network for real-world unstable data to fill the gap. D$^3$R tackles the drift via decomposition and reconstruction. In the decomposition procedure, we utilize data-time mix-attention to dynamically decompose long-period multivariate time series, overcoming the limitation of the local sliding window. The information bottleneck is critical yet difficult to determine in the reconstruction procedure. To avoid retraining once the bottleneck changes, we control it externally by noise diffusion and directly reconstruct the polluted data. The whole model can be trained end-to-end. Extensive experiments on various real-world datasets demonstrate that D$^3$R significantly outperforms existing methods, with a 11% average relative improvement over the previous SOTA models. Chengsen Wang, Zirui Zhuang, Qi Qi 0001, Jingyu Wang 0001, Haifeng Sun 0001, Jianxin Liao |
NeurIPS | 7 |
| 2023 | Poster: PipeLLM: Pipeline LLM Inference on Heterogeneous Devices with Sequence SlicingabstractLarge Language Models (LLMs) has fostered the creation of innovative requirements. Locally deployed LLMs for micro-enterprise mitigates potential issues such as privacy infringements and sluggish response. However, they are hampered by the limitations in computing capability and memory space of possessed devices. We introduce PipeLLM, which allocates the model across devices commensurate with their computing capabilities. It enables the parallel execution of layers with slicing input sequence along the token dimension. PipeLLM demonstrates the potential to accelerate LLM inference with heterogeneity devices, offering a solution for LLM deployment in micro-enterprise hardware environment. Ruilong Ma, Jingyu Wang 0001, Qi Qi 0001, Haifeng Sun 0001, Zirui Zhuang, Jianxin Liao |
SIGCOMM | 7 |
| 2023 | TADL: Fault Localization with Transformer-based Anomaly Detection for Dynamic Microservice SystemsabstractDue to the complexity of microservice architecture, it is difficult to accomplish efficient microservice anomaly detection and localization tasks and achieve the target of high system reliability. For rapid failure recovery and user satisfaction, it is significant to detect and locate anomalies fast and accurately in microservice systems. In this paper, we propose an anomaly detection and localization model based on Transformer, named TADL (Transformer-based Anomaly Detector and Locator), which models the temporal features and dynamically captures container relationships using Transformer with sandwich structure. TADL uses readily available container performance metrics, making it easy to implement in already-running container clusters. Evaluations are conducted on a sock-shop dataset collected from a real microservice system and a publicly available dataset SMD. Empirical studies on the above two datasets demonstrate that TADL can outperform baseline methods in the performance of anomaly detection, the latency of anomaly detection, and the effect of anomalous container localization, which indicates that TADL is useful in maintaining complex and dynamic microservice systems in the real world. Yuewei Li, Jingyu Wang 0001, Qi Qi 0001, Jing Wang 0039, Jianxin Liao |
SANER | 7 |
| 2023 | Unsupervised Portrait Drawing Generation for Free StylesabstractArtistic portrait drawing (APDrawing) generation has seen progress in recent years. However, due to the naturally high scarcity and artistry, it is difficult to collect large‐scale labeled and paired data and generally divide drawing styles into several specific recognized categories. Existing works suffer from the limited labeled data and naive manual division of drawing styles according to the corresponding artists. They cannot adapt to the actual situations, for example, a single artist might have multiple drawing styles and APDrawings from different artists might share similar styles. In this paper, we propose to use unlabeled and unpaired data and perform the task in an unsupervised manner. Without manual division of drawing styles, we take each portrait drawing as a unique style and introduce self‐supervised feature learning to learn free styles for unlabeled portrait drawings. Besides, we devise a style bank and a decoupled cycle structure to take over two main considerations in the task: generation quality and style control. Extensive experiments show that our model is more adaptable to different style inputs than state‐of‐the‐art methods. Jianxin Liao, Jingyu Wang 0001, Qi Qi 0001, Haifeng Sun 0001, Zirui Zhuang, Cong Liu 0046 |
Int. J. Intell. Syst. | 2 |
| 2023 | Semantic-enhanced Contrastive Learning for Session-based Recommendation
Yulong Wang 0001, Tongcun Liu, Lei Zhang 0094, Wei Li 0119, Jianxin Liao |
Knowl. Based Syst. | 6 |
| 2023 | Pose-Guided Hierarchical Graph Reasoning for 3-D Hand Pose Estimation From a Single Depth ImageabstractEstimating 3-D hand pose estimation from a single depth image is important for human-computer interaction. Although depth-based 3-D hand pose estimation has made great progress in recent years, it is still difficult to deal with some complex scenes, especially the issues of serious self-occlusion and high self-similarity of fingers. Inspired by the fact that multipart context is critical to alleviate ambiguity, and constraint relations contained in the hand structure are important for the robust estimation, we attempt to explicitly model the correlations between different hand parts. In this article, we propose a pose-guided hierarchical graph convolution (PHG) module, which is embedded into the pixelwise regression framework to enhance the convolutional feature maps by exploring the complex dependencies between different hand parts. Specifically, the PHG module first extracts hierarchical fine-grained node features under the guidance of hand pose and then uses graph convolution to perform hierarchical message passing between nodes according to the hand structure. Finally, the enhanced node features are used to generate dynamic convolution kernels to generate hierarchical structure-aware feature maps. Our method achieves state-of-the-art performance or comparable performance with the state-of-the-art methods on five 3-D hand pose datasets: 1) HANDS 2019; 2) HANDS 2017; 3) NYU; 4) ICVL; and 5) MSRA. Pengfei Ren 0001, Haifeng Sun 0001, Jiachang Hao, Qi Qi 0001, Jingyu Wang 0001, Jianxin Liao |
IEEE Trans. Cybern. | 6 |
| 2023 | Exploiting Spatial-Temporal Behavior Patterns for Fraud Detection in Telecom NetworksabstractFraud detection in telecom network is a crucial problem that threatens users’ privacy and property security. In recent years, fraudsters adopt more advanced camouflage strategies to avoid being detected by traditional algorithms. To deal with these new types of fraud, it is necessary to analyze the integrated spatial-temporal features, which are rarely involved in existing literature. In this article, we propose a novel fraud detection model based on the intertwined spatial-temporal patterns of user behaviors. Specifically, we first introduce the extension of statistical and interactive features to dynamic call patterns, and build a probabilistic model to simulate users’ call behaviors. Then the sequential patterns reflecting users’ own behaviors are obtained by the mixture Hidden Markov Models, and the structural patterns reflecting the collaboration between users in the telecom network are obtained by the attention-based Graph-SAGE model. Finally, our model outputs a fraud score for each user to detect potential fraudsters. We conduct extensive experiments on a real-world telecom dataset. The experimental results demonstrate that our intertwined spatial-temporal call patterns can effectively represent user behavior and improve the accuracy of fraud detection compared with state-of-the-art methods. The results also validate the efficiency and the interpretability of our model. Guojun Chu, Jingyu Wang 0001, Qi Qi 0001, Haifeng Sun 0001, Shimin Tao, Hao Yang 0006, Jianxin Liao, Zhu Han 0001 |
IEEE Trans. Dependable Secur. Comput. | 7 |
| 2023 | RTHop: Real-Time Hop-by-Hop Mobile Network Routing by Decentralized Learning With Semantic AttentionabstractMulti-access Edge Computing and ubiquitous smart devices help serve end-users efficiently by providing emerging edge-deployed services. On the other hand, more heavy and time-varying traffic loads are generated in mobile edge networks, so that an efficient traffic forwarding mechanism is highly required to handle the routing problem in complex and highly dynamic edge environments. Thus, Deep Reinforcement Learning (DRL) is introduced since it can work in a model-free approach. However, previous centralized DRL-based methods work in a turn-based way that mismatches the real-time property of routing. In this paper, we propose a real-time and distributed learning approach, RTHop, to adapt to the volatile environment and realize a hop-by-hop routing. The Multi-Agent Deep Reinforcement Learning (MADRL) and the Real-Time Markov Decision Process (RTMDP) are used to alleviate network congestion and maximize the utilization of network resources. By joining with the self-attention mechanism, RTHop obtains the semantics from elements of the network state to help agents learn the importance of each element on routing. Experiment results show that RTHop not only overcomes the weakness of conventional turn-based DRL methods but also achieves the increase of delivered packet ratios and effective throughput compared with other routing methods. Bo He 0003, Jingyu Wang 0001, Qi Qi 0001, Haifeng Sun 0001, Jianxin Liao |
IEEE Trans. Mob. Comput. | 5 |
| 2023 | Timely and Accurate Bitrate Switching in HTTP Adaptive Streaming With Date-Driven I-Frame PredictionabstractIn today's Internet, bandwidth dynamics are inevitable, and hence, the bitrate for live streaming applications should also be dynamically adjusted. However, in existing HTTP-based adaptive streaming (HAS), bitrate switching can only be performed at segment boundaries, making decisions unresponsive and often inaccurate. In this paper, we start from a close investigation on the impact of the segment length in HAS and accordingly presentVHAS, an extension towards intelligent variable-length segmentation, which makes client-side decisions based on the massive amount of real-time information from the network and viewers. VHAS implements a smart trigger mechanism that balances accuracy and overhead for variable-length segmentation. We further develop an adaptive bitrate switching algorithm with data-driven I-frame prediction, which is tailored to individual viewers to minimize bitrate mismatches. We evaluate VHAS via extensive trace-driven simulations, and our results demonstrate that compared with state-of-the-art solutions, VHAS achieves 15%–49% gains in QoE, with a noticeable bandwidth reduction of 37%–57%. Tongtong Feng, Qi Qi 0001, Jingyu Wang 0001, Jianxin Liao, Jiangchuan Liu |
IEEE Trans. Multim. | 4 |
| 2023 | Fine-Grained Text-to-Video Temporal Grounding from Coarse BoundaryabstractText-to-video temporal grounding aims to locate a target video moment that semantically corresponds to the given sentence query in an untrimmed video. In this task, fully supervised works require text descriptions for each event along with its temporal segment coordinate for training, which is labor-consuming. Existing weakly supervised works require only video-sentence pairs but cannot achieve satisfactory performance. However, many available annotations in the form of coarse temporal boundaries for sentences are ignored and unexploited. These coarse boundaries are common in streaming media platform and can be collected in a mechanical manner. We propose a novel approach to perform fine-grained text-to-video temporal grounding from these coarse boundaries. We take dense video captioning as base task and leverage the trained captioning model to identify the relevance of each video frame to the sentence query according to the frame participation in event captioning. To quantify the frame participation in event captioning, we proposeevent activation sequence, a simple method that highlights the temporal regions which have high correlations to the text modality in videos. Experiments on modified ActivityNet Captions and a use case demonstrate the promising fine-grained performance of our approach. Jiachang Hao, Haifeng Sun 0001, Pengfei Ren 0001, Jingyu Wang 0001, Qi Qi 0001, Jianxin Liao |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2023 | Standing on the Shoulders of Giants: Cross-Slice Federated Meta Learning for Resource Orchestration to Cold-Start SliceabstractNetwork slicing is a key technology in 6G communication systems to support numerous vertical applications for all scenes while providing resources on demand. Due to more time-varying and dynamic traffic flows, it is difficult for traditional methods to manage complex and highly dynamic 6G networks. Therefore, intelligent method such as Deep Reinforcement Learning (DRL) is employed into network management since DRL is a model-free and experience-driven approach. However, it is difficult to leverage one DRL model to provide customized intra-slice orchestration for various applications because of their diverse flow characteristics and service requirements. Moreover, training a DRL model is notoriously time-consuming so that it is not realistic for network operator to individually orchestrate customized network slice for each application with the DRL algorithm. Additionally, the data privacy of each application should be considered into the training process. In this paper, we propose a Federated Meta Reinforcement Learning (FedMRL) approach to tackle the cold-start problem in network slice orchestration, while reserving the data privacy. The Meta Reinforcement Learning (MRL) is leveraged to train a meta policy for rapidly learn a local policy for a specific slice orchestration task by finding a common initialization that allows for a quick adaptation towards each optimal solution. With the help of federated learning setting, the training process of meta policy is not required to collect raw data of applications to the centralized server. Experimental results show that FedMRL outperforms three baselines in terms of overall costs, end-to-end latency and convergence speed. Tianjian Dong, Qi Qi 0001, Jingyu Wang 0001, Zirui Zhuang, Haifeng Sun 0001, Jianxin Liao, Zhu Han 0001 |
IEEE/ACM Trans. Netw. | 6 |
| 2023 | Multi-SP Network Slicing Parallel Relieving Edge Network ConflictabstractNetwork slicing is rapidly prevailing in the edge network, which provides computing, network, and storage resources for various services. When the multiple service providers (SPs) respond to their tenants in parallel, individual decisions on the dynamic and shared edge network may lead to resource conflicts, which affects the delivery of network slicing services. Existing works ignore resource interaction and coordination in the multi-SP scenario, which is not in line with the actual situation. Indeed, the complexity of resource interaction caused by the coexistence of multiple SP policies increases the difficulty to solve the formulated optimization model. In this article, we focus on the multi-SP network slicing deployment in parallel. The coordination of network resources between SPs is designed as an effective multi-agent communication mechanism that is merged into multi-agent deep reinforcement learning (MADRL). To deal with dynamic edge networks, we design the neurons hotplugging learning which realizes scalability without a high cost of model retraining. Experiments on real and random networks demonstrate that the proposed multi-SP network slicing mechanism can successfully learn coordination policies and easily adapt to various network scales. It improves the accepted requests by 7.4%, reduces resource conflicts by 14.5%, and shortens the model convergence time by 83.3%. Rongxin Han, Dezhi Chen, Song Guo 0001, Jingyu Wang 0001, Qi Qi 0001, Lu Lu 0015, Jianxin Liao |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2023 | Identifying Users Across Social Media Networks for Interpretable Fine-Grained Neighborhood Matching by Adaptive GATabstractThe primary concern of numerous online social media network (SMN) platforms is how to provide users with effective and personalized web services. To achieve this goal, SMN platforms typically begin by collecting user preferences based on user behaviors (e.g., browsing history, posts) or user profiles. However, the effective information about a specific user on a single SMN platform is limited and monotonous, preventing a comprehensive reflection of the user's preferences. Therefore, recognizing anonymous but identical users across two SMNs to integrate their information is crucial for enhancing web services. Clearly, cross-platform research has the potential to aid in the resolution of numerous problems in service computing theory and applications. Therefore, in this article, we present theCross-PlatformUserMatcher (CPUM) framework, which attempts to map users into a union vector space and then performs user matching based on distance metrics. In particular, we introduce a GNN-based encoderAdaptiveGraphAttention Network (AdaGAT) for modeling user attributes and topology jointly in the social networks to capture two typical alignment principles: topology consistency and attribute consistency. Moreover, we derive AdaGAT from the heuristic of the spectral network alignment technique FINAL, which theoretically guarantees AdaGAT's efficacy. To the best of our knowledge, AdaGAT is the first representation-based alignment model to integrate these two alignment principles synergistically. In addition, two position encoding schemes are introduced to prevent alignment confusion that commonly arises with GNN-based alignment models. Extensive experiments on real-world datasets validate the superiority of the proposed framework. Wei Tang 0013, Haifeng Sun 0001, Jingyu Wang 0001, Cong Liu 0046, Qi Qi 0001, Jing Wang 0039, Jianxin Liao |
IEEE Trans. Serv. Comput. | 7 |
| 2022 | Probability Correlation Learning for Anomaly Detection based on Distribution-Constrained AutoencoderabstractNetwork anomaly detection provides a reliable and stable service to detect faults and prevent security attacks effectively. However, existing detection methods still encounter many challenges. The supervised learning method is unsuitable because the anomaly samples are seriously sparse and hard to label. Unsupervised learning, as a promising method, is widely used while the discriminative features are ignored when reconstructing from the normal feature space. This paper proposes a novel probability correlation learning based on autoencoder called PCDetect, a semi-supervised learning method. Since we assumed the anomaly samples deviate from the distribution of normal samples, approximating the distribution of original data is proposed as an efficient preprocessing methodology to capture the discriminative features. Moreover, an encoder-decoder neural network associated with the proposed loss function is designed to learn the low-dimensional feature representation from raw data and constrain the latent representation to follow different referenced distributions based on a few anomaly labels. In this way, The correlation of the referenced distribution and the reconstruction of latent representation will be used to quantify the probability of anomaly. Extensive experiments are conducted on two public real-world datasets, NSL-KDD and UNSW-NB15. Results show the proposed PCDetect can efficiently cope with the imbalance and high-dimensional issues compared with several popular supervised learning methods, significantly improve the accuracy, and reduce the false rate as a whole compared with unsupervised learning. Jihua Wu, Lei Zhang 0094, Cong Liu 0046, Qi Qi 0001, Jingyu Wang 0001, Tong Xu 0002, Jianxin Liao |
APNOMS | 7 |
| 2022 | Modeling Aspect Correlation for Aspect-based Sentiment Analysis via Recurrent Inverse Learning GuidanceabstractAspect-based sentiment analysis (ABSA) aims to distinguish sentiment polarity of every specific aspect in a given sentence. Previous researches have realized the importance of interactive learning with context and aspects. However, these methods are ill-studied to learn complex sentence with multiple aspects due to overlapped polarity feature. And they do not consider the correlation between aspects to distinguish overlapped feature. In order to solve this problem, we propose a new method called Recurrent Inverse Learning Guided Network (RILGNet). Our RILGNet has two points to improve the modeling of aspect correlation and the selecting of aspect feature. First, we use Recurrent Mechanism to improve the joint representation of aspects, which enhances the aspect correlation modeling iteratively. Second, we propose Inverse Learning Guidance to improve the selection of aspect feature by considering aspect correlation, which provides more useful information to determine polarity. Experimental results on SemEval 2014 Datasets demonstrate the effectiveness of RILGNet, and we further prove that RILGNet is state-of-the-art method in multiaspect scenarios. Longfeng Li, Haifeng Sun 0001, Qi Qi 0001, Jingyu Wang 0001, Jing Wang 0039, Jianxin Liao |
COLING | 6 |
| 2022 | Mining Multi-View Information: A Strong Self-Supervised Framework for Depth-based 3D Hand Pose and Mesh EstimationabstractIn this work, we study the cross-view information fusion problem in the task of self-supervised 3D hand pose estimation from the depth image. Previous methods usually adopt a hand-crafted rule to generate pseudo labels from multi-view estimations in order to supervise the network training in each view. However, these methods ignore the rich semantic information in each view and ignore the complex dependencies between different regions of different views. To solve these problems, we propose a cross-view fusion network to fully exploit and adaptively aggregate multi-view information. We encode diverse semantic information in each view into multiple compact nodes. Then, we introduce the graph convolution to model the complex dependencies between nodes and perform cross-view information interaction. Based on the cross-view fusion network, we propose a strong self-supervised framework for 3D hand pose and hand mesh estimation. Furthermore, we propose a pseudo multi-view training strategy to extend our framework to a more general scenario in which only single-view training data is used. Results on NYU dataset demonstrate that our method outperforms the previous self-supervised methods by 17.5% and 30.3% in multi-view and single-view scenarios. Meanwhile, our framework achieves comparable re-sults to several strongly supervised methods. Pengfei Ren 0001, Haifeng Sun 0001, Jiachang Hao, Jingyu Wang 0001, Qi Qi 0001, Jianxin Liao |
CVPR | 6 |
| 2022 | Can Shuffling Video Benefit Temporal Bias Problem: A Novel Training Framework for Temporal Grounding
Jiachang Hao, Haifeng Sun 0001, Pengfei Ren 0001, Jingyu Wang 0001, Qi Qi 0001, Jianxin Liao |
ECCV (36) | 6 |
| 2022 | A Temporal Consistency Enhancement Algorithm Based on Pixel Flicker Correction
Junfeng Meng, Qiwei Shen 0001, Yangliu He, Jianxin Liao |
ICONIP (4) | 4 |
| 2022 | Parallel Network Slicing for Multi-SP ServicesabstractNetwork slicing is rapidly prevailing in edge cloud, which provides computing, network and storage resources for various services. When the multiple service providers (SPs) respond to their tenants in parallel, individual decisions on the dynamic and shared edge cloud may lead to resource conflicts. The resource conflicts problem can be formulated as a multi-objective constrained optimization model; however, it is challenging to solve it due to the complexity of resource interactions caused by co-existing multi-SP policies. Therefore, we propose a CommDRL scheme based on multi-agent deep reinforcement learning (MADRL) and multi-agent communication to tackle the challenge. CommDRL can coordinate network resources between SPs with less overhead. Moreover, we design the neurons hotplugging learning in CommDRL to deal with dynamic edge cloud, which realizes scalability without a high cost of model retraining. Experiments demonstrate that CommDRL can successfully obtain deployment policies and easily adapt to various network scales. It improves the accepted requests by 7.4%, reduces resource conflicts by 14.5%, and shortens the model convergence time by 83.3%. Rongxin Han, Dezhi Chen, Song Guo 0001, Xiaoyuan Fu, Jingyu Wang 0001, Qi Qi 0001, Jianxin Liao |
ICPP | 7 |
| 2022 | FedNKD: A Dependable Federated Learning Using Fine-tuned Random Noise and Knowledge DistillationabstractMultimedia retrieval models need the ability to extract useful information from large-scale data for clients. As an important part of multimedia retrieval, image classification model directly affects the efficiency and effect of multimedia retrieval. We need a lot of data to train a image classification model applied to multimedia retrieval task. However, with the protection of data privacy, the data used to train the model often needs to be kept on the client side. Federated learning is proposed to use data from all clients to train one model while protecting privacy. When federated learning is applied, the distribution of data across different clients varies greatly. Disregarding this problem yields a final model with unstable performance. To enable federated learning to work dependably in the real world with complex data environments, we propose FedNKD, which utilizes knowledge distillation and random noise. The superior knowledge of each client is distilled into a central server to mitigate the instablity caused by Non-IID data. Importantly, a synthetic dataset is created by some random noise through back propagation of neural networks. The synthetic dataset will contain the abstract features of the real data. Then we will use this synthetic dataset to realize the knowledge distillation while protecting users' privacy. In our experimental scenarios, FedNKD outperforms existing representative algorithms by about 1.5% in accuracy. Shaoxiong Zhu, Qi Qi 0001, Zirui Zhuang, Jingyu Wang 0001, Haifeng Sun 0001, Jianxin Liao |
ICMR | 6 |
| 2022 | Query-aware video encoder for video moment retrieval
Jiachang Hao, Haifeng Sun 0001, Pengfei Ren 0001, Jingyu Wang 0001, Qi Qi 0001, Jianxin Liao |
Neurocomputing | 6 |
| 2022 | A New-Dynamic Adaptive Data Rate Algorithm of LoRaWAN in Harsh EnvironmentabstractThe adaptive data rate (ADR) algorithm is used in LoRaWAN, allocating an appropriate transmission rate for terminal equipment to improve channel utilization and reduce power consumption. However, the standard ADR algorithm is only suitable for static terminal equipment. In addition, due to the complexity of the external environment, the selected data rate will not match the new environment and communication packets will be lost. Therefore, this article proposed a novel and more effective ADR algorithm called new-dynamic ADR (ND-ADR). The algorithm mainly solves two problems, i.e., the standard ADR algorithm cannot be applied to mobile terminal devices and the poor communication quality and high packet loss rate in harsh environments. In this study, we also developed a frame for LoRaWAN (FLoRaWAN), a simulation framework for the star network topology LoRaWAN in OPNET. Furthermore, we built an OKUMURA–HATA model and additionally introduced a noise factor$\beta$to simulate the loss of wireless communication in harsh environments. Finally, extensive simulation results showed that the number of data packets required by different end nodes for rate allocation was different. Compared with the standard ADR algorithm, even for mobile nodes in harsh environments, the ND-ADR algorithm reduced network energy consumption by about 13%, reduced network delay by about 18%, and increased effective throughput by about 15%. Therefore, the improved ND-ADR algorithm is more suitable for wireless communication of removable nodes in harsh environments. Its advantages are better awareness of link environment, faster data rate regulation, improved channel utilization, and further reduction of network energy consumption. Xianghui Chen, Jianxin Liao, Weixian Song, Xihai Zhang |
IEEE Internet Things J. | 4 |
| 2022 | Following the Correct Direction: Renovating Sparsified SGD Towards Global Optimization in Distributed Edge LearningabstractDistributed edge learning collaborates powerful edge devices to train a shared global model. Since the frequent communication between the server and workers is very expensive, it is desired to accelerate the learning process. The gradient sparsification is an efficient method that only uploads a small subset of gradient elements. However, most existing works neglect the distributed nature of local datasets, and consequently the local gradients uploaded by edge devices cannot follow the global correct optimization direction well, which results in the loss of accuracy. In this paper, we propose a new gradient sparsification with a renovating mechanism, called Global Renovating Stochastic Gradient Descent (GRSGD). GRSGD utilizes the previous-round global gradient to estimate the current global one and renovates the current zero-sparsified local gradients. It mitigates the communication overhead while making the convergence direction of training closer to the global optimization, accelerating the distributed edge learning process. We provide a theoretical convergence guarantee for our algorithm based on the non-convex assumption, which better fits most deep learning problems. With extensive experiments in PyTorch, we show that GRSGD effectively accelerates the learning process with a smaller communication cost and a faster convergence rate on most training tasks. For example, on ImageNet MnasNet, GRSGD cuts down the gradient size from 8.47MB to 2.13MB while achieving 9.6%+ higher accuracy. Wanyi Ning, Haifeng Sun 0001, Xiaoyuan Fu, Qi Qi 0001, Jingyu Wang 0001, Jianxin Liao, Zhu Han 0001 |
IEEE J. Sel. Areas Commun. | 7 |
| 2022 | Extractive Dialogue Summarization Without Annotation Based on Distantly Supervised Machine Reading Comprehension in Customer ServiceabstractGiven a long dialogue, the dialogue summarization system aims to obtain a shorter highlight which retains the important information in the original text. For the customer service scenarios, the summaries of most dialogues between an agent and a user focus on several fixed key points, such as users’ question, users’ purpose, the agent’s solution, and so on. Traditional extractive methods are difficult to extract all predefined key points exactly. Furthermore, there is a lack of large-scale and high-quality extractive summarization datasets containing the annotation for key points. Moreover, the speaker’s role information is ignored or not fully utilized in previous work. In order to solve the above challenges, we propose a Distant Supervision based Machine Reading Comprehension model for extractive Summarization (DSMRC-S). DSMRC-S transforms the summarization task into the machine reading comprehension problem, to fetch key points from the original text exactly according to the predefined questions. In addition, a distant supervision method is proposed to alleviate the lack of eligible extractive summarization datasets. What’s more, a speaker’s role token and the solver classification task are proposed to make full use of speaker’s role information. We conduct experiments on a real-world summarization dataset collected in customer service scenarios, and the results show that the proposed method outperforms the strong baseline methods by 6 percentage points on ROUGE$_L$. Haifeng Sun 0001, Jingyu Wang 0001, Qi Qi 0001, Jianxin Liao |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2022 | Towards Intelligent Provisioning of Virtualized Network Functions in Cloud of Things: A Deep Reinforcement Learning Based ApproachabstractCloud of Things (CoT) is an integration of Internet of Things (IoT) and cloud computing, where Network Function Virtualization (NFV) can dynamically provide Virtualized Network Functions (VNFs) for IoT devices based on service-specific requirements. The provisioning of VNFs in CoT is formulated as an online decision-making problem, but widely used methods mostly focus on characterizing the environment using simple models to obtain the optimal solution. Valuable historical experience on provisioning for the best long-term benefits is ignored and Quality of Service (QoS) requirements for different types of CoT services are also not considered, which leads to inefficient and coarse-gained provisioning. In this article, an intelligent provisioning framework of VNFs is proposed for adaptive CoT resource scheduling according to traffic identification of heterogeneous network services. The framework leverages a Deep Reinforcement Learning (DRL)-based model to make decisions based on the complexity of network environments and traffic variances. In this model, a policy gradient DRL algorithm, namely, Policy Optimization using Kronecker-Factored Trust Region (POKTR) is adopted to obtain the stable performance by a novel surrogate objective function. Experimental results verify that our framework improves the QoS in CoT by real-time VNFs provisioning. The DRL-based model with POKTR algorithm reduces network congestion and achieves higher throughput than other DRL algorithms. Bo He 0003, Jingyu Wang 0001, Qi Qi 0001, Haifeng Sun 0001, Jianxin Liao |
IEEE Trans. Cloud Comput. | 5 |
| 2022 | GSSP: Eliminating Stragglers Through Grouping Synchronous for Distributed Deep Learning in Heterogeneous ClusterabstractDistributed deep learning has been widely used in training deep neural networks, especially for big models on massive datasets. Parameter Server (PS) architecture is the most popular distributed training framework, which can flexibly design the global parameter update manner. However, when scaling to complex heterogeneous clusters, stragglers make it difficult for existing distributed paradigms on PS framework to balance between synchronous waiting and staleness, which slows down the model training sharply. In this article, we proposeGrouping Stale Synchronous Parallel(GSSP) scheme, which groups workers with similar performance together. Group servers coordinate intra-group workers using Stale Synchronous Parallel while they communicate with each other asynchronously to eliminate stragglers and refine the model weights. We further proposeGrouping Dynamic Tok-K Sparsification(GDTopK), which dynamically adjusts the upload ratio for each group so as to make communication volume differentiated and mitigate inter-group iteration speed gap. We have conducted experiments on LeNet-5 on MNIST, ResNet-18, VGG-19 on Cifar-10, and Seq2Seq on Multi30k. Results show that GSSP accelerates the training by 46%$\sim$120%, with less than 1 percent accuracy drop. And GDTopK can make up for part of the lost accuracy. Haifeng Sun 0001, Zhiyi Gui, Song Guo 0001, Qi Qi 0001, Jingyu Wang 0001, Jianxin Liao |
IEEE Trans. Cloud Comput. | 6 |
| 2022 | DLI-Net: Dual Local Interaction Network for Fine-Grained Sketch-Based Image RetrievalabstractFine-grained sketch-based image retrieval (FG-SBIR) is considered an ideal method of image retrieval due to the rich and easily accessible characteristics of sketches. It aims to find the most similar photo from the photo gallery based on the input sketch. Most previous works follow the paradigm that extracting global feature first and then projecting the features of sketch and photo to unified embedding feature space using triplet loss. However, the global feature is not appropriate for extracting the crucial fine-grained information. Based on this principle, we propose a Dual Local Interaction Network (DLI-Net). DLI-Net explores an effective and efficient way to utilize local features for FG-SBIR. Specifically, we first propose a Local Feature Extractor to extract mid-level local features. Then, in response to the problems brought by local features, we propose a Dual Interaction Module, which contains Self Interaction Module and Cross Interaction Module. Self Interaction Module speeds up retrieval by eliminating the redundant local features of background. Cross Interaction Module solves the spatial misalignment by making the sketches interact with photos. Extensive experiments on six commonly used datasets show that our DLI-Net outperforms state-of-the-art competitors by a significant margin with a reasonable retrieval speed. Moreover, to the best of our knowledge, DLI-Net is the first model that beats humans on all six datasets. Besides, DLI-Net also performs best on cross-category fine-grained sketch-based image retrieval task, which further demonstrates local features are more appropriate for FG-SBIR. Haifeng Sun 0001, Jingyu Wang 0001, Qi Qi 0001, Ce Ge, Jianxin Liao |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2022 | Exploring Local Detail Perception for Scene Sketch Semantic SegmentationabstractIn this paper, we aim to explore the fine-grained perception ability of deep models for the newly proposed scene sketch semantic segmentation task. Scene sketches are abstract drawings containing multiple related objects. It plays a vital role in daily communication and human-computer interaction. The study has only recently started due to a main obstacle of the absence of large-scale datasets. The currently available dataset SketchyScene is composed of clip art-style edge maps, which lacks abstractness and diversity. To drive further research, we contribute two new large-scale datasets based on real hand-drawn object sketches. A general automatic scene sketch synthesis process is developed to assist with new dataset composition. Furthermore, we propose to enhancing local detail perception in deep models to realize accurate stroke-oriented scene sketch segmentation. Due to the inherent differences between hand-drawn sketches and natural images, extreme low-level local features of strokes are incorporated to improve detail discrimination. Stroke masks are also integrated into model training to guide the learning attention. Extensive experiments are conducted on three large-scale scene sketch datasets. Our method achieves state-of-the-art performance under four evaluation metrics and yields meaningful interpretability via visual analytics. Ce Ge, Haifeng Sun 0001, Yi-Zhe Song, Zhanyu Ma, Jianxin Liao |
IEEE Trans. Image Process. | 5 |
| 2022 | A Dual-Branch Self-Boosting Framework for Self-Supervised 3D Hand Pose EstimationabstractAlthough 3D hand pose estimation has made significant progress in recent years with the development of the deep neural network, most learning-based methods require a large amount of labeled data that is time-consuming to collect. In this paper, we propose a dual-branch self-boosting framework for self-supervised 3D hand pose estimation from depth images. First, we adopt a simple yet effective image-to-image translation technology to generate realistic depth images from synthetic data for network pre-training. Second, we propose a dual-branch network to perform 3D hand model estimation and pixel-wise pose estimation in a decoupled way. Through a part-aware model-fitting loss, the network can be updated according to the fine-grained differences between the hand model and the unlabeled real image. Through an inter-branch loss, the two complementary branches can boost each other continuously during self-supervised learning. Furthermore, we adopt a refinement stage to better utilize the prior structure information in the estimated hand model for a more accurate and robust estimation. Our method outperforms previous self-supervised methods by a large margin without using paired multi-view images and achieves comparable results to strongly supervised methods. Besides, by adopting our regenerated pose annotations, the performance of the skeleton-based gesture recognition is significantly improved. Pengfei Ren 0001, Haifeng Sun 0001, Jiachang Hao, Qi Qi 0001, Jingyu Wang 0001, Jianxin Liao |
IEEE Trans. Image Process. | 6 |
| 2021 | Question-Driven Span Labeling Model for Aspect-Opinion Pair ExtractionabstractAspect term extraction and opinion word extraction are two fundamental subtasks of aspect-based sentiment analysis. The internal relationship between aspect terms and opinion words is typically ignored, and information for the decision-making of buyers and sellers is insufficient. In this paper, we explore an aspect–opinion pair extraction (AOPE) task and propose a Question-Driven Span Labeling (QDSL) model to extract all the aspect–opinion pairs from user-generated reviews. Specifically, we divide the AOPE task into aspect term extraction (ATE) and aspect-specified opinion extraction (ASOE) subtasks; we first extract all the candidate aspect terms and then the corresponding opinion words given the aspect term. Unlike existing approaches that use the BIO-based tagging scheme for extraction, the QDSL model adopts a span-based tagging scheme and builds a question–answer-based machine-reading comprehension task for an effective aspect–opinion pair extraction. Extensive experiments conducted on three tasks (ATE, ASOE, and AOPE) on four benchmark datasets demonstrate that the proposed method significantly outperforms state-of-the-art approaches. Yulong Wang 0001, Tongcun Liu, Jingyu Wang 0001, Lei Zhang 0094, Jianxin Liao |
AAAI | 6 |
| 2021 | Spatial Temporal Enhanced Contrastive and Pretext Learning for Skeleton-based Action RepresentationabstractIn this paper, we focus on unsupervised representation learning for skeleton-based action recognition. The critical issue of this task is extracting discriminative spatial-temporal information from skeleton sequences to form action representation. To better solve this, we propose a novel unsupervised framework named contrastive-pretext spatial-temporal network (CP-STN), aiming to achieve accurate action recognition by better exploiting discriminative spatial-temporal enhanced features from massive unlabeled data. We combine contrastive and pretext tasks learning paradigms in one framework by using asymmetric spatial and temporal augmentations to enable network extracting discriminative representations with spatial-temporal information fully. Furthermore, graph-based convolution is used as the backbone to explore natural spatial-temporal graph information in skeleton data. Extensive experimental results show that our CP-STN significantly boosts the performance of existing skeleton-based action representations learning networks and achieves state-of-the-art accuracy on two challenging benchmarks in both unsupervised and semi-supervised settings. Yiwen Zhan 0001, Pengfei Ren 0001, Haifeng Sun 0001, Jingyu Wang 0001, Qi Qi 0001, Jianxin Liao |
ACML | 7 |
| 2021 | Joint-Aware Regression: Rethinking Regression-Based Method for 3D Hand Pose Estimation
Xiaozheng Zheng, Pengfei Ren 0001, Haifeng Sun 0001, Jingyu Wang 0001, Qi Qi 0001, Jianxin Liao |
BMVC | 6 |
| 2021 | Towards Efficient Inference: Adaptively Cooperate in Heterogeneous IoT Edge ClusterabstractNew applications such as smart homes, autonomous vehicles are leading an increasing research topic of convolutional neural network (CNN) based inference on IoT edge devices. Unfortunately, this scenario meets a huge roadblock caused by the limited computing resources owned by these devices. One popular solution is to execute inference on an edge cluster with parallelization schemes instead of on a single device. However, the heterogeneous edge devices and varied neural layers bring challenges to this process. In this paper, we propose a pipelined cooperation scheme (PICO) to efficiently execute CNN inference for edge devices. Our goal is to maximize throughput by reducing redundant computing meanwhile to keep the inference latency under a certain value. PICO divides the neural layers and edge devices into several stages. The input data is fed into the first stage and the inference result is produced at the last stage. These stages compose an inference pipeline. The execution time of each stage is optimized to approach the maximum throughput as close as possible. We also implement an adaptive framework to choose the best inference scheme under different workloads. In our experiment with 8 RaspberryPi devices, the average inference latency can be reduced by$1.7\sim 6.5\times$under different workloads, and the throughput can be improved by$1.8\sim 6.2\times$under various network settings. Qi Qi 0001, Jingyu Wang 0001, Song Guo 0001, Jianxin Liao |
ICDCS | 5 |
| 2021 | Prefix-Graph: A Versatile Log Parsing Approach Merging Prefix Tree with Probabilistic GraphabstractLogs play an important part in analyzing system behavior and diagnosing system failures. As the basic step of log analysis, log parsing converts raw log messages into structured log templates. However, existing log parsing approaches are not adaptive and versatile enough to ensure their high accuracy on all types of datasets. In particular, it is required to design regular expressions or fine-tune the hyper-parameters manually for the best performance. In this paper, we propose Prefix-Graph, an online versatile log parsing approach. Prefix-Graph is a probabilistic graph structure extended from prefix tree. It iteratively merges together two branches which have high similarity in probability distribution, and represents log templates as the combination of cut-edges in root-to-leaf paths of the graph. Since no domain knowledge is used and all the parameters are fixed, Prefix-Graph can be easily applied to different log datasets without any additional manual work. We evaluate our approach on 10 real-world datasets and 117GB log messages obtained from Huawei. The experimental results demonstrate that Prefix-Graph achieves the highest average accuracy of 0.975 and the smallest standard deviation of 0.037. Our approach is superior to baseline methods in terms of adaptability and versatility. Guojun Chu, Jingyu Wang 0001, Qi Qi 0001, Haifeng Sun 0001, Shimin Tao, Jianxin Liao |
ICDE | 6 |
| 2021 | UMDSF: Unified Model With Dynamic-Static Features for Personalized RecommendationabstractTypically, existing works utilize static methods to extract the latent feature representation of user and item reviews, neglecting the time signals and behavior patterns hidden in the user-item interaction history, which may fail to capture users' instant interests and items' temporal attributes. Moreover, there is no framework that unifies recent behavior sequences and reviews. Therefore, in this paper, we first define dynamic and static features to describe users' short- and long-term preferences and items' temporal and inherent attributes. We then design feature extractors to capture these latent factors simultaneously from recent behavior sequences and reviews. Then, we propose a novel unified framework to extract and fuse these fine-grained characteristics, named unified model with dynamic-static features (UMDSF). Specifically, the proposed model extracts both temporal sequence and review features by two parallel feature extractors based on self-attention and a multi-head attention mechanism. Subsequently, an adaptive fusion module is utilized to combine the fine-grained representations for the downstream recommendation tasks. Extensive experiments on four real-world datasets demonstrate the superiority of UMDSF and additional ablation studies verify the effectiveness of the components designed in the proposed model. Siyuan Lou, Yulong Wang 0001, Tongcun Liu, Jianxin Liao |
IJCNN | 4 |
| 2021 | Accelerating DNN Inference by Edge-Cloud CollaborationabstractDeep neural networks (DNN) have become indispensable tools for intelligent applications today. The demand for deploying DNN on the edge devices increases dramatically. Unfortunately, it is challenging because the DNN inference is computation-intensive, but edge devices are always resource-constraint. Prior solutions attempted to address these challenges with collaboration between cloud and edge devices, but they do not take the inference request rate into account. However, the inference delay will increase dramatically while the request rate becomes higher.In this paper, we propose a scheme to dynamic partition DNN into two or three parts and distribute them at the edge and cloud, achieving the lowest delay with the change of request rate. The scheme selects the optimal partition points of DNN with a layer evaluation model (LEM) and a total delay prediction model (DPM) under different request rates. The experiments of distributed deploying AlexNet, VGG, NiN and ResNet DNN models on image classification dataset ImageNet show that the proposed scheme significantly reduces the total end-to-end latency by fully using both the edge and cloud resources. It reduces the inference delay by 1.3 to 1.6 times and improves the throughput 1.2 to 1.7 times compared to the state of art partition approach. Jianan Chen 0010, Qi Qi 0001, Jingyu Wang 0001, Haifeng Sun 0001, Jianxin Liao |
IPCCC | 5 |
| 2021 | SAR: Spatial-Aware Regression for 3D Hand Pose and Mesh Reconstruction from a Monocular RGB Imageabstract3D hand reconstruction is a popular research topic in recent years, which has great potential for VR/AR applications. However, due to the limited computational resource of VR/AR equipment, the reconstruction algorithm must balance accuracy and efficiency to make the users have a good experience. Nevertheless, current methods are not doing well in balancing accuracy and efficiency. Therefore, this paper proposes a novel framework that can achieve a fast and accurate 3D hand reconstruction. Our framework relies on three essential modules, including spatial-aware initial graph building (SAIGB), graph convolutional network (GCN) based belief maps regression (GBBMR), and pose-guided refinement (PGR). At first, given image feature maps extracted by convolutional neural networks, SAIGB builds a spatial-aware and compact initial feature graph. Each node in this graph represents a vertex of the mesh and has vertex-specific spatial information that is helpful for accurate and efficient regression. After that, GBBMR first utilizes adaptive-GCN to introduce interactions between vertices to capture short-range and long-range dependencies between vertices efficiently and flexibly. Then, it maps vertices’ features to belief maps that can model the uncertainty of predictions for more accurate predictions. Finally, we apply PGR to compress the redundant vertices’ belief maps to compact-joints’ belief maps with the pose guidance and use these joints’ belief maps to refine previous predictions better to obtain more accurate and robust reconstruction results. Our method achieves state-of-the-art performance on four public benchmarks, FreiHAND, HO-3D, RHD, and STB. Moreover, our method can run at a speed of two to three times that of previous state-of-the-art methods. Our code is available at https://github.com/zxz267/SAR. Xiaozheng Zheng, Pengfei Ren 0001, Haifeng Sun 0001, Jingyu Wang 0001, Qi Qi 0001, Jianxin Liao |
ISMAR | 6 |
| 2021 | Context-Aware Anomaly Detection in Attributed Networks
Jianxin Liao, Jingyu Wang 0001, Qi Qi 0001, Haifeng Sun 0001 |
KSEM | 2 |
| 2021 | Efficient Depth Completion Network Based on Dynamic Gated Fusion
Zhengyang Mu, Qi Qi 0001, Jingyu Wang 0001, Haifeng Sun 0001, Jianxin Liao |
KSEM | 5 |
| 2021 | DLA-Net for FG-SBIR: Dynamic Local Aligned Network for Fine-Grained Sketch-Based Image RetrievalabstractFine-grained sketch-based image retrieval is considered as an ideal alternative to keyword-based image retrieval and image search by image due to the rich and easily accessible characteristics of sketches. Previous works always follow a paradigm that first extracting image global feature with convolution neural network and then optimizing the model with triplet loss. Many efforts on narrowing the domain gap and extracting discriminating features are made by these works. However, they ignored that the global feature is not good at capturing fine-grained details. In this paper, we emphasize the local features are more discriminating than global feature in FG-SBIR and explore an effective way to utilize local features. Specifically, Local Aligned Network (LA-Net) is proposed first, which solves FG-SBIR by directly aligning the mid-level local features. Experiment manifests it can beat all previous baselines and is easy to implement. LA-Net is hoped to be a new strong baseline for FG-SBIR. Next, Dynamic Local Aligned Network (DLA-Net) is proposed to enhance LA-Net. The question of spatial misalignment caused by the abstraction of the sketch is not considered by LA-Net. To solve this question, a dynamic alignment mechanism is introduced into LA-Net. This new mechanism makes the sketch interact with the photo and dynamically decide where to align according to the different photos. The Experiment indicates DLA-Net successfully addresses the question of spatial misalignment. It gains a significant performance boost over LA-Net and outperforms the state-of-the-art in FG-SBIR. To the best of our knowledge, DLA-Net is the first model that beats humans on all datasets---QMUL FG-SBIR, QMUL Handbag, and Sketchy. Haifeng Sun 0001, Qi Qi 0001, Jingyu Wang 0001, Ce Ge, Lejian Zhang, Jianxin Liao |
ACM Multimedia | 7 |
| 2021 | Few-Shot Class-Adaptive Anomaly Detection with Model-Agnostic Meta-LearningabstractAnomaly detection in encrypted traffic is a growing problem, and many approaches have been proposed to solve it. However, those approaches need to be trained in the massive of normal traffic and specific-class abnormal traffic, so to achieve good results in that specific-class. For a new anomaly class with few labeled samples, the effectiveness of existing approaches will decline sharply. How to train a model using only a few anomaly samples to detect unseen new anomaly classes in training is a huge challenge. In this paper, we propose a Few-shot Class-adaptive Anomaly Detection framework (FCAD) with model-agnostic meta-learning (MAML) to meet this challenge. Given an input network flow, FCAD first extracts statistical features by feature extractor and feature selector, and time-series features using LSTM-based AutoEncoder. Then, FCAD designs a MAML-based few-shot anomaly detection model, relying on the episodic training paradigm and learning from the collection of K-way-M-shot classification tasks, which can mimic the few-shot regime faced at test time during training. Finally, FCAD uses the pre-trained model to adapt the new class by a few iterations steps. Our goal is to detect anomaly traffic in a before unseen anomaly class with only a few samples. A reliable solution to few-shot anomaly detection will have huge potential for real-world applications since it is expensive and arduous to collect a massive amount of data onto the new anomaly class; extensive experimental results demonstrate the effectiveness of our proposed approach. Tongtong Feng, Qi Qi 0001, Jingyu Wang 0001, Jianxin Liao |
Networking | 4 |
| 2021 | Distant Supervision based Machine Reading Comprehension for Extractive Summarization in Customer ServiceabstractGiven a long text, the summarization system aims to obtain a shorter highlight while keeping important information on the original text. For customer service, the summaries of most dialogues between an agent and a user focus on several fixed key points, such as user's question, user's purpose, the agent's solution, and so on. Traditional extractive methods are difficult to extract all predefined key points exactly. Furthermore, there is a lack of large-scale and high-quality extractive summarization datasets containing key points. In order to solve the above challenges, we propose a Distant Supervision based Machine Reading Comprehension model for extractive Summarization (DSMRC-S). DSMRC-S transforms the summarization task into the machine reading comprehension problem, to fetch key points from the original text exactly according to the predefined questions. In addition, a distant supervision method is proposed to alleviate the lack of eligible extractive summarization datasets. We conduct experiments on a large-scale summarization dataset collected in customer service scenarios, and the results show that the proposed DSMRC-S outperforms the strong baseline methods by 4 points on ROUGE-L. Cao Liu, Jingyu Wang 0001, Shujie Hu, Fan Yang 0087, Guanglu Wan, Jiansong Chen, Jianxin Liao |
SIGIR | 9 |
| 2021 | Effective Scheduler for Distributed DNN Training Based on MapReduce and GPU Cluster
Jie Xu 0003, Jingyu Wang 0001, Qi Qi 0001, Haifeng Sun 0001, Jianxin Liao |
J. Grid Comput. | 5 |
| 2021 | Spatial-aware stacked regression network for real-time 3D hand pose estimation
Pengfei Ren 0001, Haifeng Sun 0001, Weiting Huang, Jiachang Hao, Daixuan Cheng, Qi Qi 0001, Jingyu Wang 0001, Jianxin Liao |
Neurocomputing | 8 |
| 2021 | Mean Field Deep Reinforcement Learning for Fair and Efficient UAV ControlabstractUnmanned aerial vehicles (UAVs) can provide flexible network coverage services. UAVs can be applied in a large number of scenarios, such as emergency communication and network access in areas without terrestrial network coverage. However, UAVs are limited to relatively short communication range and restricted energy resources. In extreme conditions such as disasters, there may also be a problem that the communication bandwidth is limited and the UAV cannot communicate with the server with a large amount of information, so a decentralized solution is expected. In addition, the interaction between multiple objectives and multiple UAVs leads to a huge state space, which makes large-scale practical applications difficult. To simplify complex interactions, we modeled the UAV control problem with mean-field game (MFG). We propose a new UAV control method, the mean-field trust region policy optimization (MFTRPO), which uses the MFG method to construct the Hamilton-Jacobi-Bellman/Fokker-Planck-Kolmogorov equation that obtains the optimal solution and solves the difficulties in the practical application through the trust region policy optimization and neural network feature embedding methods. The proposed method: 1) maximizes communication efficiency while ensuring fair communication range and network connectivity; 2) fuses the mean-field theory with deep reinforcement learning techniques; and 3) is scalable and adaptive. We conduct extensive simulations for performance evaluation. The simulation results have shown that MFTRPO significantly and consistently outperforms two commonly used baseline methods in terms of coverage, fairness, and energy consumption. Dezhi Chen, Qi Qi 0001, Zirui Zhuang, Jingyu Wang 0001, Jianxin Liao, Zhu Han 0001 |
IEEE Internet Things J. | 5 |
| 2021 | Adaptive and Robust Routing With Lyapunov-Based Deep RL in MEC Networks Enabled by BlockchainsabstractThe most recent development of the Internet of Things brings massive timely sensitive and bursty data flows. Also, joint optimization on storage, computation, and communication is in need for multiaccess edge computing frameworks. The adaptive network control has been explored using deep reinforcement learning (RL), but it is not sufficient for bursty network traffic flows, especially when the network traffic pattern may change over time. We formulate the routing control in an environment with time-variant link delays as a Lyapunov optimization problem. We identify that there is a tradeoff between optimization performance and modeling accuracy when the propagation delays are included. We propose a novel deep RL (DRL)-based adaptive network routing method to tackle the issues mentioned above. A Lyapunov optimization technique is used to reduce the upper bound of the Lyapunov drift, improving queuing stability in networked systems. By modeling the network traffic pattern using the Markovian arrival process, we show that network routing problems can be modeled as Markov decision processes and value-iteration-based RL methods can be used to solve them. We design a blockchain-based protocol using proof of elapsed time consensus mechanism to ensure a trustworthy network statistics information exchange for the routing framework. Experiment results show that the proposed method can learn a routing policy and adapt to the changing environment. The proposed method outperforms the baseline backpressure method in multiple settings and converges faster than existing methods. Moreover, the DRL module can effectively learn a better estimation of the long-term Lyapunov drift and penalty functions, providing superior results in terms of the backlog size, end-to-end latency, age of information, and throughput. Furthermore, the blockchain-based network statistics exchange can provide the routing framework against malicious nodes. In addition, the proposed model performs well under various topologies, and thus can be used in general cases. Zirui Zhuang, Jingyu Wang 0001, Qi Qi 0001, Jianxin Liao, Zhu Han 0001 |
IEEE Internet Things J. | 4 |
| 2021 | Pattern and content controlled response generation
Haifeng Sun 0001, Daixuan Cheng, Jingyu Wang 0001, Qi Qi 0001, Jianxin Liao |
Inf. Process. Manag. | 5 |
| 2021 | Network-aware task selection to reduce multi-application makespan in cloud
Jie Xu 0039, Jingyu Wang 0001, Qi Qi 0001, Jianxin Liao, Haifeng Sun 0001, Zhu Han 0001, Tonghong Li |
J. Netw. Comput. Appl. | 4 |
| 2021 | Generative Adversarial Network-Based Transfer Reinforcement Learning for Routing With Prior KnowledgeabstractWith the incremental deployment of software defined networking, the routing algorithms have gained more power on observability and controllability. Deep reinforcement learning, as an experience-driven approach, shows considerable potential in routing problem with the help of the centralized controller. It is an adaptive, lightweight, and model-free approach to coping with dynamic runtime status, large-scale traffic, and heterogeneous objective of SDN routing. However, it is still not suitable for the variable and complex emerging networks, because the huge training cost prevents fast convergence in a varying or discrepant environment. In this paper, we propose a transfer reinforcement learning algorithm to improve the training efficiency, and handle the variation in network status and topology. Specifically, we leverage the generative adversarial network to learn domain-invariant features that is suitable for deep reinforcement learning-based routing in different network environments. This mechanism utilizes the previous model and accelerates the training process. We implement our routing algorithm in the production level software switches and controller, while evaluating it comprehensively with many topologies and network status distributions. The experimental results show that our work not only outperforms the state-of-the-art deep reinforcement learning-based routing frameworks, but also has more training efficiency than the naive transfer learning algorithm both on different topologies and network status distributions. Tianjian Dong, Qi Qi 0001, Jingyu Wang 0001, Alex X. Liu, Haifeng Sun 0001, Zirui Zhuang, Jianxin Liao |
IEEE Trans. Netw. Serv. Manag. | 7 |
| 2021 | DeepCC: Multi-Agent Deep Reinforcement Learning Congestion Control for Multi-Path TCP Based on Self-AttentionabstractWith the development of the Internet of Things (IoT) and 5G, there are ubiquitous smart devices and network functions providing emerging network services efficiently and optimally through building many network connections based on WiFi, LTE/5G, Ethernet, and etc. The Multipath TCP (MPTCP) protocol that enables these devices to establish multiple paths for simultaneous data transmission, has been a widely used extension of standard TCP in smart devices and network functions. On the other hand, more heavy and time-varying traffic loads are generated in an MPTCP network, so that an efficient congestion control mechanism that schedules the traffic between multiple subflows and avoids congestion is highly required. In this paper, we propose a decentralized learning approach, DeepCC, to adapt to the volatile environments and realize the efficient congestion control. The Multi-Agent Deep Reinforcement Learning (MADRL) is used to learn a policy of congestion control for each subflow according to the real-time network states. To deal with the problem of the fixed state space and slow convergence, we adopt two self-attention mechanisms to receive the states and train the policy, respectively. Due to the asynchronous design of DeepCC, the learning process will not introduce extra delay and overhead on the decision-making process. Experiment results show that DeepCC consistently outperforms the well-known heuristic method and DRL-based MPTCP congestion control method in terms of goodput and jitter. Besides, DeepCC with the attention mechanism reduces convergence time by about 50% and increase goodput by about 80% compared with the commonly used structures of neural networks. Bo He 0003, Jingyu Wang 0001, Qi Qi 0001, Haifeng Sun 0001, Jianxin Liao, Chunning Du, Alex X. Liu, Zhu Han 0001 |
IEEE Trans. Netw. Serv. Manag. | 5 |
| 2021 | An integrated model based on deep multimodal and rank learning for point-of-interest recommendation
Jianxin Liao, Tongcun Liu, Hongzhi Yin, Tong Chen 0005, Jingyu Wang 0001, Yulong Wang 0001 |
World Wide Web | 1 |
| 2020 | Adversarial and Domain-Aware BERT for Cross-Domain Sentiment AnalysisabstractCross-domain sentiment classification aims to address the lack of massive amounts of labeled data.It demands to predict sentiment polarity on a target domain utilizing a classifier learned from a source domain.In this paper, we investigate how to efficiently apply the pre-training language model BERT on the unsupervised domain adaptation.Due to the pre-training task and corpus, BERT is taskagnostic, which lacks domain awareness and can not distinguish the characteristic of source and target domain when transferring knowledge.To tackle these problems, we design a post-training procedure, which contains the target domain masked language model task and a novel domain-distinguish pre-training task.The post-training procedure will encourage BERT to be domain-aware and distill the domain-specific features in a self-supervised way.Based on this, we could then conduct the adversarial training to derive the enhanced domain-invariant features.Extensive experiments on Amazon dataset show that our model outperforms state-of-the-art methods by a large margin.The ablation study demonstrates that the remarkable improvement is not only from BERT but also from our method. Chunning Du, Haifeng Sun 0001, Jingyu Wang 0001, Qi Qi 0001, Jianxin Liao |
ACL | 5 |
| 2020 | Dual Attention-Based Adversarial Autoencoder for Attributed Network EmbeddingabstractExisting embedding methods for Attributed Network aim to learn low-dimensional embeddings for nodes, which can preserve both consistency and complementarity for network structures and node attributes. The main assumption is that nodes with similar structures and/or similar attributes should be close in the embedding space. In reality, nodes with similar attributes might be far away from each other in topology and vice versa. The conflict is often caused by noisy links or incomplete network structures. Previous methods either independently project embeddings based on the assumption without considering the conflicts, or encode embeddings into a shared space ignoring the complementarity. In this paper, we propose a Dual Attention-based Adversarial Attributed Network Embedding framework (DAANE) to preserve the consistency and complementarity between structures and attributes, and reduce the conflict caused by their discrepancy. DAANE includes an attribute attention mechanism designed to detect and weakening the impact of noisy links and a structure attention mechanism applied to assign weights to network structures of different scales and capture a more complete global context. Furthermore, we develop efficient adversarial learning when combining the two heterogeneous embeddings. The adversarial auto-encoder projects embeddings of attributes and structures into the same space. Meanwhile, it completely circumvents the interference of various types of noise by removing the constraints of embedding space. Extensive experiments on three realworld network datasets indicate that the proposed model achieves state-of-the-art results. Jianxin Liao, Jingyu Wang 0001, Qi Qi 0001, Haifeng Sun 0001 |
ECAI | 2 |
| 2020 | Semi-Supervised Sentence Classification Based on User Polarity in the Social ScenariosabstractThe data sparsity is the main challenge in sentence classification in social scenarios, the recent methods incorporate user information by encoding user node in the user-relation network to alleviate this issue. However, the connection between users is not always available due to privacy protection or other commercial reasons. Thus, in this paper, a concept called user polarity is proposed to quantify the tendency of sentences published by a user which are categorized into the same class. Then a self-training framework based on user polarity is proposed, which incorporates user information without connection between users, to alleviate the data sparsity in sentence classification. A regularization term is used to strengthen the prediction of the model in some special points, and a sample selector is designed to reduce the noise in the pseudo-labeled data generated in self-training process. Besides, some hard samples are selected to improve the retraining process. The experimental results conducted on SemEval 2019 task 8 indicate that our method performs significantly better than other three semi-supervised methods and achieves state-of-the-art performance on this benchmark. Haifeng Sun 0001, Jingyu Wang 0001, Qi Qi 0001, Jianxin Liao |
ICASSP | 5 |
| 2020 | DeepHop on Edge: Hop-by-hop Routing byDistributed Learning with Semantic AttentionabstractMulti-access Edge Computing (MEC) and ubiquitous smart devices help serve end-users efficiently and optimally through providing emerging edge-deployed services. Meanwhile, heavy and time-varying traffic loads are produced in the edge network, so that an efficient traffic forwarding mechanism is required. In this paper, we propose a parallel and distributed learning approach, DeepHop, to adapt to the volatile environments and realize hop-by-hop routing. The Multi-Agent Deep Reinforcement Learning (MADRL) is used to alleviate the edge network congestion and maximize the utilization of network resources. DeepHop determines the routing among edge network nodes for heterogeneous types of traffic according to the current workload and capability. By joining with an attention mechanism, DeepHop obtains the semantics from the elements of the network state to help the agents learn the importance of each element on routing. Experiment results show that DeepHop achieves the increase of successfully transmitted packets by 15% compared with the state-of-the-art algorithms. Besides, DeepHop with an attention mechanism reduces convergence time by nearly half compared with the common-used structures of neural networks. Bo He 0003, Jingyu Wang 0001, Qi Qi 0001, Haifeng Sun 0001, Zirui Zhuang, Cong Liu 0046, Jianxin Liao |
ICPP | 7 |
| 2020 | Adaptive and Robust Network Routing Based on Deep Reinforcement Learning with Lyapunov OptimizationabstractThe most recent development of the Internet of Things brings massive timely-sensitive and yet bursty data flows. The adaptive network control has been explored using deep reinforcement learning, but it is not sufficient for extremely bursty network traffic flows, especially when the network traffic pattern may change over time. We model the routing control in an environment with time-variant link delays as a Lyapunov optimization problem. We identify that there is a tradeoff between optimization performance and modeling accuracy when the propagation delays are included. We propose a novel deep reinforcement learning-based adaptive network routing method to tackle the issues mentioned above. A Lyapunov optimization technique is used to reduce the upper bound of the Lyapunov drift, which leads to improved queuing stability in networked systems. Experiment results show that the proposed method can learn a routing control policy and adapt to the changing environment. The proposed method outperforms the baseline backpressure method in multiple settings, and converges faster than existing methods. Moreover, the deep reinforcement learning module can effectively learn a better estimation of the longterm Lyapunov drift and penalty functions, and thus it provides superior results in terms of the backlog size, end-to-end latency, age of information, and throughput. Extensive experiments also show that the proposed model performs well under various topologies, and thus the proposed model can be used in general cases. Also the user can adjust the preference parameter at ant time without the need to retrain the neural networks. Zirui Zhuang, Jingyu Wang 0001, Qi Qi 0001, Jianxin Liao, Zhu Han 0001 |
IWQoS | 4 |
| 2020 | Learning chinese word embeddings from character structural information
Qi Qi 0001, Jianxin Liao, Haifeng Sun 0001, Jingyu Wang 0001 |
Comput. Speech Lang. | 3 |
| 2020 | Exploiting geographical-temporal awareness attention for next point-of-interest recommendation
Tongcun Liu, Jianxin Liao, Zhigen Wu, Yulong Wang 0001, Jingyu Wang 0001 |
Neurocomputing | 2 |
| 2020 | Toward Communication-Efficient Federated Learning in the Internet of Things With Edge ComputingabstractFederated learning is an emerging concept that trains the machine learning models with the local distributed data sets, without sending the raw data to the data center. But, in the Internet of Things (IoT) where the wireless network resource is constrained, the key problem of federated learning is the communication overhead for parameter synchronization, which wastes bandwidth, increases training time, and even impacts the model accuracy. Gradient sparsification has received increasing attention, which only updates significant gradients and accumulates insignificant gradients locally. However, how to preserve the accuracy after a high ratio sparsification has been ignored in the literature. In this article, a general gradient sparsification (GGS) framework is proposed for adaptive optimizers, to correct the sparse gradient update process. It consists of two important mechanisms: 1) gradient correction and 2) batch normalization (BN) update with local gradients. With gradient correction, the optimizer can properly treat the accumulated insignificant gradients, which makes the model converge better. Furthermore, updating the BN layer with local gradients can relieve the impact of delayed gradients without increasing the communication overhead. We have conducted experiments on LeNet-5, CifarNet, DenseNet-121, and AlexNet with adaptive optimizers. Results show that when 99.9% gradients are sparsified, validation data sets are maintained with top-1 accuracy. Haifeng Sun 0001, F. Richard Yu, Qi Qi 0001, Jingyu Wang 0001, Jianxin Liao |
IEEE Internet Things J. | 6 |
| 2020 | Matrix Capsule Convolutional Projection for Deep Feature LearningabstractCapsule projection network (CapProNet) has shown its ability to obtain semantic information, and spatial structural information from the raw images. However, the vector capsule of CapProNet has limitations in representing semantic information due to ignoring local information. Besides, the number of trainable parameters also increases greatly with the dimension of the feature vector. To that end, we propose a matrix capsule convolution projection (MCCP) module by replacing the feature vector with a feature matrix, of which each column represents a local feature. The feature matrix is then convoluted by columns into capsule subspaces to decrease the number of trainable parameters effectively. Furthermore, the CapDetNet is designed to explore the structural information encoding of the MCCP module based on object detection task. Experimental results demonstrate that the proposed MCCP outperforms the baselines in image classification, and CapDetNet achieves the 2.3% performance gain in object detection. Canqun Xiang, Zhennan Wang 0001, Shishun Tian, Jianxin Liao, Wenbin Zou, Chen Xu 0004 |
IEEE Signal Process. Lett. | 4 |
| 2020 | A Novel Multi-Task Learning Framework for Semi-Supervised Semantic ParsingabstractWhile sequence-to-sequence (seq2seq) models on semantic parsing have demonstrated significant performance, the need for large amounts of labeled data still hinders the application of this technology to resource-poor domains. In this work, we work on alleviating data scarcity in semantic parsing. We propose a semi-supervised semantic parsing methods by exploiting unlabeled natural utterances in a novel multi-task learning framework. Two strategies are proposed. The first one takes entity sequences as training targets to improve the representations of encoder and reduce entity-mistakes in prediction. The second one extends Mean Teacher to seq2seq model and generates more target-side data to improve the generalizability of decoder network. Experiments demonstrate that our proposed methods significantly outperform the supervised baseline and achieve more impressive improvement than previous methods. Qi Qi 0001, Haifeng Sun 0001, Jingyu Wang 0001, Xiao Liang 0005, Jianxin Liao |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2020 | Vabis: Video Adaptation Bitrate System for Time-Critical Live StreamingabstractWith the rise of time-critical and interactive scenarios, ultra-low latency has become the most urgent requirement. Adaptive bitrate (ABR) schemes have been widely used in reducing latency for live streaming services. However, the traditional solutions suffer from a key limitation: they only utilize coarse-grained chunk to solve the I-frame misalignment problem in different bitrate switching process at the cost of increasing latency. As a result, existing schemes are difficult to guarantee the timeliness and granularity of control in essence. In this paper, we use a frame-based approach to solve the I-frame misalignment problem and propose a video adaptation bitrate system (Vabis) in units of the frame for time-critical live streaming to obtain the optimal quality of experience (QoE). On the server-side, a Few-Wait ABR algorithm based on Reinforcement Learning (RL) is designed to adaptively select the bitrate of future frames by state information that can be observed, which can subtly solve the problem of I-frame misalignment. A rule-based ABR algorithm is designed to optimize the Vabis system for the weak network. On the client-side, three delay control mechanisms are designed to achieve frame-based fine-grained control. We construct a trace-driven simulator and the real live platform to evaluate the comprehensive live streaming performance. The results show that Vabis is significantly better than the existing methods with decreases in an average delay of 32%-77% and improvements in average QoE of 28-67%. Tongtong Feng, Haifeng Sun 0001, Qi Qi 0001, Jingyu Wang 0001, Jianxin Liao |
IEEE Trans. Multim. | 5 |
| 2020 | Dynamic Service Function Chain Embedding for NFV-Enabled IoT: A Deep Reinforcement Learning ApproachabstractThe Internet of things (IoT) is becoming more and more flexible and economical with the advancement in information and communication technologies. However, IoT networks will be ultra-dense with the explosive growth of IoT devices. Network function virtualization (NFV) emerges to provide flexible network frameworks and efficient resource management for the performance of IoT networks. In NFV-enabled IoT infrastructure, service function chain (SFC) is an ordered combination of virtual network functions (VNFs) that are related to each other based on the logic of IoT applications. However, the embedding process of SFC to IoT networks is becoming a big challenge due to the dynamic nature of IoT networks and the abundance of IoT terminals. In this paper, we decompose the complex VNFs into smaller virtual network function components (VNFCs) to make more effective decisions since VNF nodes and IoT network devices are usually heterogeneous. In addition, a deep reinforcement learning (DRL) based scheme with experience replay and target network is proposed as a solution that can efficiently handle complex and dynamic SFC embedding scenarios in IoT. Our simulations consider different types of IoT network topologies. The simulation results present the efficiency of the proposed dynamic SFC embedding scheme. Xiaoyuan Fu, F. Richard Yu, Jingyu Wang 0001, Qi Qi 0001, Jianxin Liao |
IEEE Trans. Wirel. Commun. | 5 |
| 2019 | Multi-task Deep Reinforcement Learning for Scalable Parallel Task SchedulingabstractThe rapid development of artificial intelligence in multiple scenarios, including machine learning, image recognition, and autonomous driving, has led to an explosion of computation jobs. These jobs are often divided into parallel child tasks and executed in distributed clusters with limited computing resources, making parallel task scheduling one of the most important research topics nowadays. Most studies about parallel task scheduling focused on formulating special scenarios and service requirements as optimization problems. However, the complicated and dynamic parallel computing environments are hard to model, predict and control, making those previous methods unscalable and unable to reflect the real scenarios. In this paper, a Multi-task Deep reinforcement learning approach for scalable parallel Task Scheduling (MDTS) is firstly devised. Generally, Deep Reinforcement Learning (DRL) is a model-free optimization algorithm for long-term control by leveraging experience, but it suffers the curse of dimensionality for decision when coping with complex parallel computing environments and jobs with diverse properties. We extend the action selection in DRL to a multi-task decision, where the output branches of multitask learning are fine-matched to parallel scheduling tasks. Child tasks of a job are accordingly assigned to distributed nodes without any human knowledge while the resource competition among parallel tasks is leveraged through shared neural network layers. Extensive experiments show that the MDTS significantly reduces the job execution time compared with least-connection scheduling and particle swarm optimization algorithm by 15.3% and 39.8% respectively. Moreover, MDTS outperforms the raw DRL algorithm on job execution time, load imbalance value, and total cost by 42.8%, 47.5%, and 59.0%. Lingxin Zhang, Qi Qi 0001, Jingyu Wang 0001, Haifeng Sun 0001, Jianxin Liao |
IEEE BigData | 5 |
| 2019 | Investigating Capsule Network and Semantic Feature on Hyperplanes for Text ClassificationabstractChunning Du, Haifeng Sun, Jingyu Wang, Qi Qi, Jianxin Liao, Chun Wang, Bing Ma. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Chunning Du, Haifeng Sun 0001, Jingyu Wang 0001, Qi Qi 0001, Jianxin Liao |
EMNLP/IJCNLP (1) | 5 |
| 2019 | Capsule Network with Interactive Attention for Aspect-Level Sentiment ClassificationabstractChunning Du, Haifeng Sun, Jingyu Wang, Qi Qi, Jianxin Liao, Tong Xu, Ming Liu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Chunning Du, Haifeng Sun 0001, Jingyu Wang 0001, Qi Qi 0001, Jianxin Liao, Tong Xu 0002 |
EMNLP/IJCNLP (1) | 5 |
| 2019 | AGRM: Attention-Based Graph Representation Model for Telecom Fraud DetectionabstractFraud detection is an increasingly important and difficult task in the modern technological environment. Existing classification methods for fraud detection are based on the call behavior attributes, such as duration and frequency of callings. With the evolution of the fraud mode, the methods based on self-attributes cannot detect the fraudsters who camouflage themselves by behaving like normal accounts. So we propose an Attention-based Graph Representation learning Model(AGRM), which takes into account the node's self-attributes and its context information. In our model, an attention architecture is introduced to learn the discriminative representation by focusing on the most informative context. On this basis, we design a self-attention mechanism to adjust the contribution of the node's self-attributes and its neighbors'. Extensive experiments show that the proposed method achieved significant accuracy improvement compared with existing fraud detection methods. Jianxin Liao, Jingyu Wang 0001, Qi Qi 0001 |
ICC | 2 |
| 2019 | ALSR: An Adaptive Label Screening and Relearning Approach for Anomaly DetectionabstractAnomaly detection using KPIs (Key Performance Indicators) is key to AIOps (Artificial Intelligence for IT Operations). Recent anomaly detection approaches have adopted machine learning to detect anomalies on the perspective of individual time points more than events. These approaches do not make effective use of the labels of continuous anomaly intervals, nor do they pay attention to the differences among anomaly points. The detection performance is therefore not high enough. In this paper, we propose an anomaly detection approach named ALSR, which uses a label screening model and a relearning model to analyze and utilize the continuous anomaly intervals of KPIs in finer granularity. The label screening algorithm takes advantage of the continuity of anomaly intervals to remove unnecessary data from the training set, so as to better suit to interval-oriented anomaly detection. The relearning algorithm reclassifies the true/false positive points within range of detected anomalies, thus effectively reduces the number of false positive points. ALSR uses statistical characteristics and time series models for feature extraction, and the feature set is proved to better describe the characteristics of KPIs. We conduct comprehensive experiments on 25 KPIs, and the total F-score of ALSR is 0.965, which outperforms state-of-the-art anomaly detection approaches. Yuhan Jing, Qi Qi 0001, Jingyu Wang 0001, Tongtong Feng, Jianxin Liao |
ISCC | 5 |
| 2019 | A Geographical-Temporal Awareness Hierarchical Attention Network for Next Point-of-Interest RecommendationabstractObtaining insight into user mobility for next point-of-interest (POI) recommendations is a vital yet challenging task in location-based social networking. Information is needed not only to estimate user preferences but to leverage sequence relationships from user check-ins. Existing approaches to understanding user mobility gloss over the check-in sequence, making it difficult to capture the subtle POI-POI connections and distinguish relevant check-ins from the irrelevant. We created a geographically-temporally awareness hierarchical attention network (GT-HAN) to resolve those issues. GT-HAN contains an extended attention network that uses a theory of geographical influence to simultaneously uncover the overall sequence dependence and the subtle POI-POI relationships. We show that the mining of subtle POI-POI relationships significantly improves the quality of next POI recommendations. A context-specific co-attention network was designed to learn changing user preferences by adaptively selecting relevant check-in activities from check-in histories, which enabled GT-HAN to distinguish degrees of user preference for different check-ins. Tests using two large-scale datasets (obtained from Foursquare and Gowalla) demonstrated the superiority of GT-HAN over existing approaches and achieved excellent results. Tongcun Liu, Jianxin Liao, Zhigen Wu, Yulong Wang 0001, Jingyu Wang 0001 |
ICMR | 2 |
| 2019 | ResidualDenseNetwork: A Simple Approach for Video Person IdentificationabstractVideo identification is an important task in the practical application and industry. Based on the iQIYI-VID-2019 dataset, ACM International Conference on Multimedia and iQIYI co-hosted the celebrity video identification challenge. We take part in the competition, propose a new feature fusion method and design a residual dense network which can improve video identification performance in the complex scenes. Only with face features, we achieve 0.9035 in mean Average Precision(mAP) which win the second place on the leadboard. At the same time, it is the best score only with official features. It is worth mention that the flops of our model is only 0.5G and the time required to predict the entire test dataset is only 2 sim 5 minutes. Our method takes accuracy and speed into account, which has a strong practical significance. Zixuan Huang 0009, Weizhao Chen, Qiwei Shen 0001, Jianxin Liao |
ACM Multimedia | 5 |
| 2019 | ALSR: An adaptive label screening and relearning approach for interval-oriented anomaly detection
Jingyu Wang 0001, Yuhan Jing, Qi Qi 0001, Tongtong Feng, Jianxin Liao |
Expert Syst. Appl. | 5 |
| 2019 | Deep supervised hashing network with integrated regularisationabstractHashing has been widely deployed to approximate nearest neighbour search for large‐scale multimedia retrieval tasks due to storage and retrieval efficiency. State‐of‐the‐art supervised hashing methods for image retrieval construct deep structures to simultaneously learn image representation and generate good hash codes, and the key step among them is simultaneously learned feature representation and binary hash code. Existing methods use similarity and regularity loss to train deep hashing systems, but these two functions usually work together but not cooperative, which may lead to inadequate performance of the whole system. In this study, a new method for training deep hashing system to learn compact binary codes is presented. The deep supervised hashing network with integrated regularisation (DSHIR) system develop the zero division restriction as a new part of the loss function, which settles the problem of cooperatively guiding the system generate similarity preserving binary codes. DSHIR system also modifies the similarity handling loss to better extract features from image data, which promotes the performance compared to existing end‐to‐end deep hashing systems. Experiments show that DSHIR yields about 10 per cent higher mean average precision on CIFAR‐10 dataset, and also promote on other evaluation indexes compared with state‐of‐the‐art systems. Jianxin Liao, Baoran Li, Jingyu Wang 0001, Qi Qi 0001, Jing Wang 0039 |
IET Image Process. | 1 |
| 2019 | Collaborative tensor-topic factorization model for personalized activity recommendation
Tongcun Liu, Jianxin Liao, Yulong Wang 0001, Jingyu Wang 0001, Qi Qi 0001 |
Multim. Tools Appl. | 2 |
| 2019 | Radio Environment Map Construction by Kriging Algorithm Based on Mobile Crowd SensingabstractIn the IoT era, 5G will enable various IoT services such as broadband access everywhere, high user and devices mobility, and connectivity of massive number of devices. Radio environment map (REM) can be applied to improve the utilization of radio resources for the access control of IoT devices by allocating them reasonable wireless spectrum resources. However, the primary problem of constructing REM is how to collect the large scale of data. Mobile crowd sensing (MCS), leveraging the smart devices carried by ordinary people to collect information, is an effective solution for collecting the radio environment information for building the REM. In this paper, we build a REM collecting prototype system based on MCS to collect the data required by the radio environment information. However, limited by the budget of the platform, it is hard to recruit enough participants to join the sensing task to collect the radio environment information. This will make the radio environment information of the sensing area incomplete, which cannot describe the radio information accuracy. Considering that the Kriging algorithm has been widely used in geostatistics principle for spatial interpolation for Kriging giving the best unbiased estimate with minimized variance, we utilize the Kriging interpolation algorithm to infer complete radio environment information from collected sample radio environment information data. The interpolation performance is analyzed based on the collected sample radio environment information data. We demonstrate experiments to analyze the Kriging interpolation algorithm interpolation results and error and compared them with the nearest neighbor (NN) and the inverse distance weighting (IDW) interpolation algorithms. Experiment results show that the Kriging algorithm can be applied to infer radio environment information data based on the collected sample data and the Kriging interpolation has the least interpolation error. Zhifeng Han, Jianxin Liao, Qi Qi 0001, Haifeng Sun 0001, Jingyu Wang 0001 |
Wirel. Commun. Mob. Comput. | 2 |
| 2018 | Fewer is More: Image Segmentation Based Weakly Supervised Object Detection with Partial Aggregation
Ce Ge, Jingyu Wang 0001, Qi Qi 0001, Haifeng Sun 0001, Jianxin Liao |
BMVC | 5 |
| 2018 | A Case-Based Decision System for Routing in Packet-Switched NetworksabstractRoute planning with global optimization objectives in graphs is a challenging task with enormous computational complexity and finding the best solution is NP-complete. In addition, the network's operational performance varies whenever the environment changes. Traditional routing schemes fail to deal with these situations. We propose a case-based decision system for routing in packet switched networks to track the networking status. We also design a graph-aware neural network to suggest and revise the solutions from the past cases. The low-level structure of the neural network is learned by fitting with the features not only from each standalone vertex but also from the neighbors of each vertex. Experiments show that the proposed system outperforms state-of-art traffic-split and traffic-engineered routing schemes. Zirui Zhuang, Jingyu Wang 0001, Qi Qi 0001, Haifeng Sun 0001, Jianxin Liao |
IPCCC | 5 |
| 2018 | ECTCP: An Explicit Centralized Congestion Avoidance for TCP in SDN-based Data CenterabstractData centers provide high aggregate bandwidth for large-scale network services. Traditional TCP protocol does not work well to meet the demand in data center. Before congestion occurs, TCP increases sending window to preempt bandwidth, resulting in buildup queues and increase latency, moreover TCP can only adjust sending window individually for each flow, which lacks global information. This mechanism is ill-suited for the diverse mix of short and long flows, and cannot provide high burst tolerance. Fortunately, Software Defined Networking (SDN) has a global view on data center network, which provides an opportunity to centralized control flows. Leveraging the feature of SDN, we propose an explicit centralized congestion avoidance mechanism (ECTCP). By gaining network topology and the information of flows, SDN controller calculates each flow's fair bandwidth, and adjusts the receive window size of each TCP connection. ECTCP does not need to modify TCP stack, which is easy to implement in the SDN-based data center. In our evaluation, ECTCP achieves fair throughput and improves fairness among different congestion control algorithms. ECTCP also reduces Flow Completion Time (FCT) for burst short flows, and keeps Round-Trip Time (RTT) in small-scale. Jiannan Bao, Jingyu Wang 0001, Qi Qi 0001, Jianxin Liao |
ISCC | 4 |
| 2018 | Game Theoretic Modeling of Interaction Problems in Datacenter Network EnvironmentsabstractCloud services have been widely deployed on Iarge-scale datacenters by Cloud Service Providers (CSPs) to provide users with better performance and reliability. However, interactions between datacenter selection and multiple traffic engineering (TE) of many physical networks operated by Internet Service Providers (ISPs) may occur, which may lead to performance loss. In this paper, we propose a cooperative approach based on Nash's bargaining solution (NBS) for CSP and ISPs to improve their performance, which considers both Pareto optimality and fairness. Then, in order to facilitate the actual deployment, we also design a decomposition algorithm to calculate the NBS. Simulations are performed to verify our analysis on different network environments. Jianxin Liao, Zilu Kang |
ISCC | 2 |
| 2018 | Graph-Aware Deep Learning Based Intelligent Routing StrategyabstractSoftware defined networking decouples the control plane and data plane, which grants more computing power for routing computations. Traditional routing methods suffer from the complex dynamics in networking, and they are facing issues such as slow convergence and performance decline. Deep learning techniques have shown preliminary results on solving the routing problem, bring more accuracy and precision compared with traditional modeling techniques. However, the deep learning architecture needs to be specially customized to learn the topological relations between switches in an efficient way. Thus, we propose a deep learning based intelligent routing strategy with revised graph-aware neural networks and we design a set of features suitable for network routing. Then we demonstrate the performance of our works by using a real-world topology and the production level software switch. The simulation result shows our work is more accurate and efficient compared to state-of-art routing strategy. Zirui Zhuang, Jingyu Wang 0001, Qi Qi 0001, Haifeng Sun 0001, Jianxin Liao |
LCN | 5 |
| 2018 | Incentive Mechanisms for Resource Scaling-out Game of Stream Big Data Analytics
Xiaoyuan Fu, Jingyu Wang 0001, Qi Qi 0001, Jianxin Liao, Tonghong Li |
J. Grid Comput. | 4 |
| 2018 | MTDE: Multi-typed data embedding in heterogeneous networks
Haifeng Sun 0001, Jianxin Liao, Jingyu Wang 0001, Qi Qi 0001 |
Neurocomputing | 2 |
| 2018 | Rapid Relevance Feedback Strategy Based on Distributed CBIR SystemabstractThis article describes the capability of online data storage which has been enhanced by the emergence of cloud datacenter development. Distributed Hash Table (DHT) based image retrieval system using locality sensitive hash (LSH) has provided an efficient way to set up distributed Content Based Image Retrieval (CBIR) frameworks. However, with the fixed LSH function adopted, LSH and other codebook-based distributed retrieval systems are facing the problem of flexibility, and also are difficult to satisfy the user's demand. In this article, LRFMIR is proposed to introduce semantic search into DHT based CBIR system. LRFMIR is established on a DHT based network, where a flexible result truncating strategy is employed to fuse provided results by using multiple features measurements. Experiments show that LRFMIR provides a higher accuracy and recall rate than single feature employed retrieval systems, and possesses good load balancing and query efficiency performance. Jianxin Liao, Baoran Li, Jingyu Wang 0001, Qi Qi 0001, Tonghong Li |
Int. J. Semantic Web Inf. Syst. | 1 |
| 2017 | Tax-Based Mechanisms for Resource Scaling-Out of Stream Big Data AnalyticsabstractCloud-based big data platforms provide physical resources for a variety of applications to analyze all forms of data. For the stream big data analytics, a participated task always needs to scale out resources when its input data increases steeply. Typically, the resource scaling out can be achieved by increasing the parallelism degree of the platform based on the experience. However, the resource scaling-out of each task produces additional cost not only from itself but also from other competitive tasks, which brings about great challenges to ensure the efficient utilization of resources. To solve this problem systematically, we consider the resource scaling-out problem as a non-cooperative game and formulate a total cost model including a risk function and a task execution time function. The total cost of resource scaling-out reflects the influence of topology structure for the benefit of a participated task. Hence, two economic classic tax-based incentive policies: Pivotal Mechanism and Externality Mechanism are applied, to stimulate the participation of tasks. We make simulations in different scenarios including node degree and different characteristics of tasks. The simulations results show that our resource scaling-out mechanism can achieve a better performance close to social optimality. Xiaoyuan Fu, Jingyu Wang 0001, Qi Qi 0001, Jianxin Liao, Tonghong Li |
PDCAT | 4 |
| 2017 | Density cluster based approach for controller placement problem in large-scale software defined networkings
Jianxin Liao, Haifeng Sun 0001, Jingyu Wang 0001, Qi Qi 0001, Kai Li 0003, Tonghong Li |
Comput. Networks | 1 |
| 2017 | A generalized design of distributed rateless codes with decreasing ripple size for multiple-access relay networks
Jianxin Liao, Lei Zhang 0094, Tonghong Li, Jingyu Wang 0001, Qi Qi 0001 |
Wirel. Networks | 1 |
| 2016 | An Approach to Improve the Cooperation between Heterogeneous SDN OverlaysabstractThe overlay network has been widely developed in recent years. There may be various overlays that co-exist with each other upon the same underlying network. These overlays have heterogeneous performance goals, and they will compete for the physical resources, so that a sub-optimal performance of the overlays may be achieved. Moreover, the heterogeneity of the overlays makes them difficult to coordinate with each other to improve their performance. We introduce the concept of SDN to the deployment of overlay network and propose an approach to make the overlays cooperate with each other. A cooperative solution is proposed for co-existing overlays to improve their performance while leveraging their heterogeneous performance goals. Simulations are performed to evaluate the cooperative solution. Ziteng Cui, Jianxin Liao, Jingyu Wang 0001, Qi Qi 0001, Jing Wang 0039 |
LCN | 2 |
| 2016 | Design of optimised multiple partial recovery LT codesabstractExisting rateless codes have a very low intermediate symbol recovery rate. Therefore, a new analysis method named iterative and small degree first (I‐SDF) is presented for the design of optimised partial recovery Luby transform codes (PR‐LTC) in this study. On the basis of I‐SDF, the required number of encoded symbols with degree d in each decoding step is calculated by an iterative optimisation algorithm. Under the proposed design, R ( R < k ) input symbols can be recovered from as few encoded symbols as possible in PR‐LTC with message length k . Furthermore, multiple PR‐LTC (M‐PR‐LTC) is proposed to recover several partial recovery point (PRPs) efficiently. The analysis process is divided into multiple stages, and the required number of encoded symbols with degree d in each decoding step is calculated by a cross‐stage iterative optimisation algorithm. In addition, the interaction of each stage is adjusted by introducing a weight for each PRP. The PR‐LTC and M‐PR‐LTC are evaluated and compared with the existing schemes. The simulation results demonstrate that PR‐LTC and M‐PR‐LTC outperform other existing schemes in terms of average overhead, average degree of encoded symbols, memory usage, bit error rate and energy consumption. Jianxin Liao, Lei Zhang 0094, Tonghong Li, Jingyu Wang 0001, Qi Qi 0001 |
IET Commun. | 1 |
| 2016 | A dual mode self-adaption handoff for multimedia services in mobile cloud computing environment
Jianxin Liao, Qi Qi 0001, Jing Wang 0039, Jingyu Wang 0001, Yufei Cao |
Multim. Tools Appl. | 1 |
| 2016 | Fusion feature for LSH-based image retrieval in a cloud datacenter
Jianxin Liao, Tonghong Li, Qi Qi 0001, Jingyu Wang 0001, Haifeng Sun 0001 |
Multim. Tools Appl. | 1 |
| 2016 | Game-theoretic model of asymmetrical multipath selection in pervasive computing environment
Jingyu Wang 0001, Jianxin Liao, Tonghong Li, Jing Wang 0039 |
Pervasive Mob. Comput. | 2 |
| 2016 | OMI-DL: An Ontology Matching FrameworkabstractThis paper focuses on matching ontologies created for similar domains through different sources. Different solutions use lexical, structural or logical processing and analysis to match ontologies. However, an important aspect is also interpreting concepts that entities are presented with and using them in relation to semantics in an ontology. The paper demonstrates analyzing and extending the concepts used to define entities in an ontology, discusses establishing and filtering matching candidates by reasoners, and then describes constructing correspondences between entities from different ontologies using lexical and semantic analysis. The experiments show that our prototype, called OMI-DL, is among the top group over many ontologies adopted from the OAEI benchmark data set. We also provide an evaluation of the OMI-DL method for matching the DOLCE+DnS Ultralite ontology and the domain ontology developed in the m:Ciudad project. Xiulei Liu, Bo Cheng 0001, Jianxin Liao, Payam M. Barnaghi, Jingyu Wang 0001 |
IEEE Trans. Serv. Comput. | 3 |
| 2015 | VecLP: A Realtime Video Recommendation System for Live TV ProgramsabstractWe propose VecLP, a novel Internet Video recommendation system working for Live TV Programs in this paper. Given little information on the live TV programs, our proposed VecLP system can effectively collect necessary information on both the programs and the subscribers as well as a large volume of related online videos, and then recommend the relevant Internet videos to the subscribers. For that, the key frames are firstly detected from the live TV programs, and then visual and textual features are extracted from these frames to enhance the understanding of the TV broadcasts. Furthermore, by utilizing the subscribers' profiles and their social relationships, a user preference model is constructed, which greatly improves the diversity of the recommendations in our system. The subscriber's browsing history is also recorded and used to make a further personalized recommendation. This work also illustrates how our proposed VecLP system makes it happen. Finally, we dispose some sort of new recommendation strategies in use at the system to meet special needs from diverse live TV programs and throw light upon how to fuse these strategies. Sheng Gao 0001, Honggang Zhang 0002, Jianxin Liao, Jun Guo 0002 |
AAAI | 6 |
| 2015 | Improving Cross-Domain Recommendation through Probabilistic Cluster-Level Latent Factor ModelabstractCross-domain recommendation has been proposed to transfer user behavior pattern by pooling together the rating data from multiple domains to alleviate the sparsity problem appearing in single rating domains. However, previous models only assume that multiple domains share a latent common rating pattern based on the user-item co-clustering. To capture diversities among different domains, we propose a novel Probabilistic Cluster-level Latent Factor (PCLF) model to improve the cross-domain recommendation performance. Experiments on several real world datasets demonstrate that our proposed model outperforms the state-of-the-art methods for the cross-domain recommendation task. Siting Ren, Sheng Gao 0001, Jianxin Liao, Jun Guo 0002 |
AAAI | 3 |
| 2015 | Interactions among Overlays and Traffic Engineering: Equilibrium and Cooperation without PaymentabstractEmerging overlay technologies have enabled to distribute content efficiently over the Internet that in some sense improves the user quality of experience. However, due to inconsistent or even conflicting objectives from the perspectives of multiple co-existing overlay networks and traffic engineering (TE), the interaction among them impacts the performance of each other and results in sub- optimum, which also has not yet been investigated in detail. In this paper, we model this interaction as an n+1-person non-cooperative game, prove the existence of Nash equilibrium point (NEP) and propose an algorithm to compute NEP. In order to overcome the inefficiency of NEP, we define a non-transferable utility (NTU) game based on Shapley NTU game theory, in which overlays and TE share cost efficiently and fairly without side payment, and we propose an algorithm to calculate the Shapley NTU value (SNTUV). Shan Jiang 0008, Jingyu Wang 0001, Jianxin Liao, Tonghong Li |
GLOBECOM | 4 |
| 2015 | Cooperative traffic management for co-existing overlaysabstractThe overlay network has been widely deployed by Service Providers to provide services. Since there are multiple SPs built upon the same ISP, their overlays are co-existing and may interfere with each other. The selfishness of overlay may lead to sub-optimal performance and traffic arrangement dilemma for overlays. To optimize the performances of overlays and maximize the benefit of SPs, we proposed a cooperative traffic management framework. Several models are applied to analyze and solve the overlay routing problem, the revenue allocation problem, and the coalition formation problem in the framework. Simulations are performed to evaluate the framework. Ziteng Cui, Jianxin Liao, Jingyu Wang 0001, Qi Qi 0001, Jing Wang 0039 |
LCN | 2 |
| 2015 | Competitive equilibrium and stable coalition in overlay environmentsabstractOverlay networks have been widely deployed upon the Internet to provide improved network services. However, the interaction between overlay and traffic engineering (TE) as well as among co-existing overlays may occur. In this paper, we adopt game theoretic approaches to analyze this hybrid interaction. Firstly, we model a situation of the hybrid interaction as an n+1-player non-cooperative game, where overlays and TE are of equal status, and prove the existence of Nash equilibrium (NE). Secondly, we model another situation of the hybrid interaction as a 1-leadern-follower Stackelberg-Nash game, where TE is the leader and coexisting overlays are followers, and prove that the cost at Stackelberg-Nash equilibrium (SNE) is at least as good as that at NE for TE. Thirdly, we propose a cooperative coalition mechanism based on Shapley value to overcome the inherent inefficiency of NE and SNE, where players can improve their performance and form stable coalitions. Shan Jiang 0008, Jianxin Liao, Jingyu Wang 0001, Tonghong Li |
LCN | 2 |
| 2015 | A coalitional game approach on improving interactions in multiple overlay environments
Jianxin Liao, Ziteng Cui, Jingyu Wang 0001, Tonghong Li, Qi Qi 0001, Jing Wang 0039 |
Comput. Networks | 1 |
| 2015 | On the collaborations of multiple selfish overlays using multi-path resources
Jingyu Wang 0001, Jianxin Liao, Tonghong Li, Jing Wang 0039 |
Peer-to-Peer Netw. Appl. | 2 |
| 2014 | Multiple features for image retrieval in distributed datacenterabstractThe emergence of cloud datacenters enhances the capability of online data storage. Since massive data is stored in datacenters, it is necessary to effectively locate interest data in such a distributed system. However, traditional search techniques only allow users to search images over exact-match keywords through a centralized index. These techniques cannot satisfy the requirements of content based image retrieval (CBIR). In this paper, we propose the scalable image retrieval framework which can efficiently support content similarity search in the distributed environment. Its key idea is to integrate image fusion features into distributed hash tables (DHTs) by exploiting the property of the locality sensitive hashing (LSH). Thus, the images with similar content are most likely gathered into the same node without the knowledge of any global information. To the best of our knowledge, there is less comprehensive study on large-scale CBIR with fusion features in the distributed environment. Jianxin Liao, Jingyu Wang 0001, Qi Qi 0001, Haifeng Sun 0001 |
APNOMS | 2 |
| 2014 | Cooperative overlay routing in a multiple overlay environmentabstractOverlay networks have been widely developed over the past few years. More and more overlays are deployed on the top of the same native network, and share the same physical resources. Competing for these physical resources, co-existing overlays may affect each other adversely. It has been showed that by using selfish overlay routing, co-existing overlays would be likely to converge to a Nash equilibrium which is sub-optimal. However, to achieve the global optimal may also cause the performance degradation of certain overlays, which make it hard to realize. Inspired by the Nash bargaining solution, a cooperative method is proposed for two co-existing overlays to achieve a near Pareto optimal. Simulations are performed to evaluate the proposed approach. The results show that the approach is effective and efficient in the multiple overlay networks environment. Ziteng Cui, Jianxin Liao, Jingyu Wang 0001, Qi Qi 0001, Jing Wang 0039 |
ICC | 2 |
| 2014 | Topology-aware Virtual Network Embedding based on multiple characteristicsabstractNetwork Virtualization provides a promising tool to allow multiple heterogeneous virtual networks to run on a shared substrate network simultaneously. A long-standing challenge in Network Virtualization is the Virtual Network Embedding (VNE) problem: how to embed virtual networks onto specific physical nodes and links in the substrate network effectively and efficiently. Recent research presents several heuristic algorithms that only consider single network topological attribute, which may lead to decreased utilization of resources. In this paper, we introduce seven complementary characteristics that reflect different topological attributes, and propose three topology-aware VNE algorithms by leveraging their respective advantages. Due to overall considering topological attributes of substrate and virtual networks through multiple characteristics, our study better coordinates node and link embedding. Extensive simulations demonstrate that our algorithms improve the long-term average revenue, acceptance ratio, and revenue/cost ratio compared to previous algorithms. Min Feng 0003, Jianxin Liao, Jingyu Wang 0001, Sude Qing, Qi Qi 0001 |
ICC | 2 |
| 2014 | Combination feature for image retrieval in the distributed datacenterabstractSince the emergence of cloud datacenters provides an enormous amount of resources easily accessible to people, it is challenging to provide an efficient search framework in such a distributed environment. However, traditional search techniques only allow users to search images over exact-match keywords through a centralized index. These methods are insufficient to meet requirements of content based image retrieval (CBIR) and more powerful search frameworks are needed. In this paper, we present LCFIR, an effective image retrieval framework for fast content location in the distributed situation. It adopts the peer-to-peer paradigm and combines color and edge features. The basic idea is to construct multiple replicas of an image's index through exploiting the property of Locality Sensitive Hashing (LSH). Thus, the indexes of similar images are probabilistically gathered into the same node without the knowledge of any global information. The empirical results show that the system is able to yield high accuracy with load balancing, and only contacts a few number of the participating nodes. Jianxin Liao, Qi Qi 0001, Jingyu Wang 0001, Haifeng Sun 0001, Shan Jiang 0008 |
ICPADS | 2 |
| 2014 | Overlay routing network construction by introducing Super-Relay nodesabstractOverlay routing has emerged as a promising approach to improve reliability and efficiency of the Internet. However, deploying a routing overlay network requires the placement and maintenance of overlay infrastructure, especially, the selection and placement of key relay nodes. Spurred by the observation that a few relay nodes with high betweenness centrality, which are called Super-Relay nodes, can provide more optimal routes for a large number of node pairs, we propose an overlay routing network constructing method by introducing Super-Relay nodes. In detail, we focus on the selection and connection of Super-Relay nodes to optimize routing quality in a resilient and scalable manner, by which we address K-Minimum Spanning Tree with Super-Relay nodes algorithm (SR-KMST). Using simulations on a real Internet topology, we conclude that our approach can provide high-quality overlay routing service, and achieve better robustness by comparing with other algorithms. Jianxin Liao, Jingyu Wang 0001, Qi Qi 0001 |
ISCC | 2 |
| 2014 | Introducing collaborations for multi-path selection of multiple selfish overlaysabstractIn complex Internet, different overlay flows are likely to share and compete the same congestible resources. We present a game-theoretic study of the selfish strategic collaboration of multiple overlays when they are allowed to use multipath transfer, which is referred as the multipath selection game. Then we consider the equilibrium in this multipath selection game model where selfish players distribute their overlay traffic. Maximization of the utility functions for each overlay is the criterion of optimality. We adopt the objective of throughput maximization to capture the most typical overlay behaviors, and use the usual TCP as the basis of our analysis. We show analytically the existence and uniqueness of Nash equilibria in these games. Furthermore, we find that the loss of efficiency of Nash equilibria can be arbitrarily large if overlays do not have resource limitations. Our simulations confirm effectiveness and TCP-friendliness of multipath transfer for a range of path number and in the presence of multiple overlay traffic. Jingyu Wang 0001, Jianxin Liao, Jing Wang 0039, Qi Qi 0001, Tonghong Li |
ISCC | 2 |
| 2014 | An image retrieval framework for distributed datacentersabstractAs massive data is stored in cloud datacenters, it is necessary to effectively locate interest data in such a distributed environment. However, since it is difficult to create a visual vocabulary due to the lack of global information, most existing systems of Content Based Image Retrieval (CBIR) only focus on global image features. In this paper, we propose a novel image retrieval framework, which efficiently incorporates the bag-of-visual-word model into Distributed Hash Tables (DHTs). Its key idea is to establish visual words for local image features by exploiting the merit of Locality Sensitive Hashing (LSH), so that similar image patches are most likely gathered into the same nodes without the knowledge of any global information. Extensive experimental results demonstrate that our approach yields high accuracy at very low cost, while keeping the load balanced. Jianxin Liao, Qi Qi 0001, Jingyu Wang 0001, Tonghong Li |
LCN | 2 |
| 2014 | A Self-Adaption Handoff Mechanism for Multimedia Services in Mobile Cloud ComputingabstractDeveloping mobile multimedia cloud services over heterogeneous wireless networks poses a challenge for service continuity. The degraded link quality and connection losses are likely to happen and these may affect service execution times and service availability in mobile cloud computing scenarios. To improve handoff quality and minimize utilized bandwidth, we propose a self-adaption handoff scheme for multimedia services in mobile cloud computing. The proposed scheme uses multipath transmission for media flows, and consists of the duplicate mode and the effective mode, which are changed according to the network condition. Analytic model and simulation are developed to investigate our new scheme. The results demonstrate that the new mechanism can realize seamless handoff for multimedia services in cloud, reduce the packet loss rate, as well as obtain a more efficient use of the scarce wireless bandwidth and the power of mobile devices. Qi Qi 0001, Jianxin Liao, Yufei Cao, Jingyu Wang 0001 |
VTC Fall | 2 |
| 2014 | Cloud service-aware location update in mobile cloud computingabstractMobile devices are becoming the primary platforms for many users who always roam around when accessing the cloud computing services. From this, the cloud computing is integrated into the mobile environment by introducing a new paradigm, mobile cloud computing. In the context of mobile computing, the battery life of mobile device is limited, and it is important to balance the mobility performance and energy consumption. Fortunately, cloud services provide both opportunities and challenges for mobility management. Taking the activities of cloud services accessing into consideration, the authors propose a service‐aware location update mechanism, which can detect the presence and location of the mobile device without traditional periodic registration update. Analytic model and simulation are developed to investigate the new mechanism. The results demonstrate that the service‐aware location update management can reduce the location update times and handoff signalling, which can efficiently save power consumption for mobile devices. Qi Qi 0001, Jianxin Liao, Yufei Cao |
IET Commun. | 2 |
| 2014 | Design of improved Luby transform codes with decreasing ripple size and feedbackabstractIn this study, the design of improved Luby transform codes with decreasing ripple size (LTC‐DRS) with feedback is presented. Under the proposed design, a new degree distribution algorithm named generalised degree distribution algorithm (GDDA) is proposed, which can achieve arbitrary ripple size revolution accurately. On the basis of GDDA, an accurate ripple size revolution based on binomial fitting is proposed, which can keep the ripple size to a suitable value throughout the decoding process. Furthermore, the authors introduce the feedback and propose a shifted ripple size revolution to diversify the degree values. The improved LTC‐DRS with feedback is evaluated and compared with the existing schemes. The simulation results demonstrate that it outperforms other existing schemes in terms of average overhead, average degree of encoded symbols, memory usage and energy consumption. Lei Zhang 0094, Jianxin Liao, Jingyu Wang 0001, Tonghong Li, Qi Qi 0001 |
IET Commun. | 2 |
| 2014 | Probe-based end-to-end overload control for networks of SIP servers
Jinzhu Wang, Jianxin Liao, Tonghong Li, Jing Wang 0039, Jingyu Wang 0001, Qi Qi 0001 |
J. Netw. Comput. Appl. | 2 |
| 2014 | Accurate sub-swarms particle swarm optimization algorithm for service composition
Jianxin Liao, Xiaomin Zhu 0002, Jingyu Wang 0001 |
J. Syst. Softw. | 1 |
| 2013 | Reducing the oscillations between overlay routing and traffic engineering by repeated game theoryabstractDue to the conflicts existing in the route objectives of overlay routing and traffic engineering, the interaction between the two selfish players may converge to an inefficient Nash equilibrium point, even if the better choice may exist. We formulate the interaction as an infinitely repeated two-player game, where overlay routing aims to minimize the average latency of overlay users and traffic engineering aims to minimize the maximum link utilization of overall network. The whole interaction process could be divided into two stages - learning stage and practice stage. The former collects the historical information and finds the best point with a simple learning algorithm, then the latter uses this point as equilibrium point and converges to it. The simulation results show that both overlay routing and traffic engineering can converge to the win-win results, and the overall network can avoid the performance volatility from endless oscillations. Jianxin Liao, Jingyu Wang 0001, Qi Qi 0001, Lei Zhang 0094 |
APCC | 2 |
| 2013 | A multi-objective service selection algorithm for service compositionabstractService composition is an efficient way to implement a service of complex business process. In service composition, service selection considering multiple QoS objectives is a challenge. Existing methods mainly utilize fitness function or constraint technique to convert multiple objectives service composition problems to single objective ones. These methods take effect with priori knowledge of problem's solution space. Besides, only one solution can be obtained in each execution of one existing method. Users can hardly acquire evenly distributed solutions with acceptable computation cost. In this paper, we propose a lightweight particle swarm optimization service selection algorithm using approximate distance and external archive mechanisms. Simulation results demonstrate that the proposed algorithm is more effective and efficient than the compared algorithm for multi-objective service composition problems. Jianxin Liao, Xiaomin Zhu 0002, Jingyu Wang 0001, Qi Qi 0001 |
APCC | 1 |
| 2013 | Accurate QoS-based service selection algorithm for service compositionabstractThe most important thing of service composition (SC) is to select optimal candidate service instances compliant with non-functional requirements (e.g. QoS and load balance constraints). Particle swarm optimization (PSO) is known as an effective and efficient algorithm, which is widely used in this process. However, the premature convergence and diversity loss of PSO may result in suboptimal solutions. In this paper, we propose an accurate sub-swarms particle swarm optimization (ASPSO) algorithm which locates optimal solutions by using sub-swarms searching grid cells in which the density of feasible solutions is high. Simulation results demonstrate that the proposed algorithm improves the accuracy of the standard PSO algorithm in service composition. Jianxin Liao, Xiaomin Zhu 0002, Jingyu Wang 0001, Qi Qi 0001 |
LCN | 1 |
| 2012 | Topology-aware virtual network embedding through bayesian network analysisabstractMultiple heterogenous virtual networks are given the ability to run on a shared infrastructure simultaneously as independent slices in the network virtualization environment. However, a major challenge is how to map multiple virtual networks, with specific node and link constraints, onto the shared substrate network, known as virtual network embedding problem. By taking topology attribute into account, topology-aware virtual network embedding algorithms efficiently improve the performance by leveraging a node ranking method based on Markov chain. However, as the basis of node ranking, the resource evaluation of node which is calculated as the product of its CPU and bandwidth may be incorrect. Moreover, a greedy matching strategy is always applied in the node mapping stage, which may lead to unnecessary bandwidth consumption by ignoring the relationships between the mapped substrate nodes and the mapping one. In this paper, we re-think the topology-aware virtual network embedding from a statistical perspective by proposing a statistical method to generate a dependency matrix representing the importance of every node and the relationships between every two nodes in the substrate network. Based on this dependency matrix, bayesian network analysis is leveraged to iteratively select the substrate node, with the closest relationship to the selected ones, to achieve node mapping process. Extensive simulations were conducted and the results show that our proposed algorithm has better performance in the long-term run. Sude Qing, Qi Qi 0001, Jingyu Wang 0001, Tong Xu 0002, Jianxin Liao |
GLOBECOM | 5 |
| 2012 | Hybrid virtual network embedding with K-core decomposition and time-oriented priorityabstractNetwork virtualization provides a powerful tool to allow multiple networks, each customized to a specific purpose, to run on a shared substrate. However, a big challenge is how to map multiple virtual networks onto specific nodes and links in the shared substrate network, known as virtual network embedding problem. Previous works in virtual network embedding can be decomposed to two classes: two-stage virtual network embedding and one-stage virtual network embedding. In this paper, by pruning the topology of virtual network using k-core decomposition, a hybrid virtual network embedding algorithm is proposed to leverage the respective advantage of the two kinds of algorithm simultaneously in the mapping process. In addition, a time-oriented scheduling policy is introduced to improve the mapping performance. We conduct extensive simulations and the results show that the proposed algorithm obtains more revenue in the long-term run. Sude Qing, Jianxin Liao, Jingyu Wang 0001, Xiaomin Zhu 0002, Qi Qi 0001 |
ICC | 2 |
| 2012 | Ontology Alignment by Combining Lexical Analysis with Consequences from ReasonersabstractAligning different ontologies from similar (or same) domains is an active field of current research. There are various solutions which process and analyze lexical, structural or semantic information to align ontologies. However, there are few solutions that focus on interpreting the concepts that entities are presented with and using them in relation to the semantics implied in an ontology. In this paper, the prototype (OACLAI) is presented to tackle this by combining lexical analysis with consequences from reasoners which reflect the semantics implied in an ontology. We evaluate OACLAI over the four real ontologies and compare it against the seven solutions. The experiments show that the accuracy of OACLAI is higher than those of others on average. Jianxin Liao, Xiulei Liu, Xiaomin Zhu 0002, Tong Xu 0002, Jingyu Wang 0001, Haifeng Sun 0001 |
ICWS | 1 |
| 2012 | Probabilistic fault localization with sliding windows
Jianxin Liao, Tonghong Li, Xiaomin Zhu 0002 |
Sci. China Inf. Sci. | 2 |
| 2012 | A distributed end-to-end overload control mechanism for networks of SIP servers
Jianxin Liao, Jinzhu Wang, Tonghong Li, Jing Wang 0039, Jingyu Wang 0001, Xiaomin Zhu 0002 |
Comput. Networks | 1 |
| 2012 | OSIA: Out-of-order Scheduling for In-order Arriving in concurrent multi-path transfer
Jingyu Wang 0001, Jianxin Liao, Tonghong Li |
J. Netw. Comput. Appl. | 2 |
| 2011 | Niching Particle Swarm Optimization Algorithm for Service CompositionabstractService composition constructs composite applications with different services to fulfill numerous service requirements. Different users have various QoS demands for composite services. In this paper, we present a service composition model considering not only multiple QoS constraints but also load balance factors. Moreover, a service selection algorithm based on niching technique and particle swarm optimization (PSO) is proposed for the service composition problem. It supports optimization problems with any kinds of constraints and objective functions. Simulation results show the proposed algorithm is effective and efficient for service composition under any circumstances. Jianxin Liao, Xiaomin Zhu 0002, Tong Xu 0002, Jingyu Wang 0001 |
GLOBECOM | 1 |
| 2011 | Introducing multipath selection for concurrent multipath transfer in the future internet
Jianxin Liao, Jingyu Wang 0001, Tonghong Li, Xiaomin Zhu 0002 |
Comput. Networks | 1 |
| 2011 | A token-bucket based notification traffic control mechanism for IMS presence service
Jianxin Liao, Jinzhu Wang, Tonghong Li, Jingyu Wang 0001, Xiaomin Zhu 0002 |
Comput. Commun. | 1 |
| 2010 | Performance Evaluation of Frequency Planning in a Novel Cellular Architecture Based on Sector RelayabstractTo avoid inter-cell interference effectively and to improve the performance of cell-edge users, a novel cellular architecture, termed Cellular Architecture Based on Sector Relay (CASR), is proposed in this paper, and two frequency planning schemes based on CASR are investigated. In CASR, fixed relays equipped with directional antennas are located at the vertex of the hexagonal cell for sectorization, thus each relay can be shared by three cells. The two frequency planning schemes are designed for different situations, one of which focuses more on the performance of cell one-hop users and the other one focuses on the performance balance between one-hop users and two-hop users. Interference analysis and performance evaluation are implemented for both one-hop users and two-hop users. System-level simulation in multi-cell and multi-user environment shows that compared with the existing cellular relay cell architecture, the proposed CASR combined with the two frequency planning schemes yield higher cell spectral efficiency. Moreover, the performance of cell-edge users is significantly improved. Lin Qu, Xiaoxiang Wang, Yulong Wang 0001, Jianxin Liao |
VTC Spring | 4 |
| 2010 | Dynamic Resource Allocation with Threshold in OFDMA-based Relay NetworksabstractIn this paper, we investigate resource allocation issue in OFDMA-based decode-and-forward cooperative networks and propose joint subcarrier and power allocation schemes. The optimal solution of this combinable allocation has high computational complexity, so we divide our solution into two steps. The first step is to distribute subcarriers to relays and destination under the assumption of equal power distribution. Here, we propose Proportional Allocation (PA) strategy to achieve tradeoff between total throughput and fairness. To further improve the system performance, we introduce threshold into PA strategy, named Proportional Allocation with Threshold (PA-T), where subcarriers with bad performance are prevented from transmitting. Next, water-filling method is adopted to distribute the power to cooperative links in order to fully utilize the limited power. Simulation results show that system performance of the proposed schemes is significantly enhanced compared with an existing resource allocation scheme. Besides, the resource allocation schemes with water-filling method notably outperform schemes with equal power allocation. Mingwei Tang, Xiaoxiang Wang, Yulong Wang 0001, Jianxin Liao |
VTC Spring | 4 |
| 2010 | Enhanced IMS Handoff Mechanism for QoS Support over Heterogeneous NetworkabstractIP multimedia subsystem (IMS) is over IP network architecture, but mobile IP cannot directly support session mobility controlled by session initiation protocol-based signaling. The long signaling delay for session reestablishment in application layer always results in session interruptions during the handoff. Therefore, handoff poses a challenge for quality of service (QoS) maintenance in IMS that targets to offer real-time multimedia applications over wireless mobile networks. The existing approaches to solve this problem depend on the advance resource reservation and the optimization of handoff control. Unfortunately, big cost of the advance resource reservation in neighboring domains is a major problem that leads to a serious signaling load and a waste of wireless bandwidth. To solve this issue, we present an enhanced IMS handoff mechanism (EHM) based on user mobility prediction to save network resources by avoiding multiple useless advance reservations. In addition, to support the heterogeneous access networks in IMS domain, EHM evolves a network selective scheme to utilize the network resources more efficiently. The architecture of EHM and the advance QoS negotiation signaling are also presented. We model the cost, the handoff delay and the session blocking probability for EHM and the previous work. Analytical and simulation results show that EHM can enhance the handoff performance, such as reducing resource reservation cost greatly, decreasing session reestablishment delay and making good use of multiple access network resources. Jianxin Liao, Qi Qi 0001, Xiaomin Zhu 0002, Yufei Cao, Tonghong Li |
Comput. J. | 1 |
| 2009 | JCCM: Joint Cluster Communities on Attribute and Relationship Data in Social Networks
Jianxin Liao, Xiaomin Zhu 0002 |
ADMA | 2 |
| 2009 | A Group Based Service Triggering Algorithm for IMS NetworkabstractTo determine the potential signaling traffic reductions, the session establishment procedures are investigated. The investigation shows that, the S-CSCF (Serving Call Session Control Function) is the major bottleneck in IMS (IP Multimedia Subsystem) and the existing 3GPP (the 3rd Generation Partnership Project) Service Triggering Algorithm (3GPP STA) increases largely the end to end session setup delay. To reduce the session setup delay and improve the system performance, a new Group based Service Triggering Algorithm (GSTA) is proposed. And then the modeling of 3GPP STA and GSTA are presented. Theoretical analysis and simulation results show that, GSTA can efficiently reduce the signaling traffic load of the S-CSCF, increase the throughput of the system and considerably reduce the session setup delay, improve IMS network quality of service. Zhaoyong Xun, Jianxin Liao, Xiaomin Zhu 0002, Yufei Cao |
ICC | 2 |
| 2008 | CDPM: Finding and Evaluating Community Structure in Social Networks
Jianxin Liao, Xiaomin Zhu 0002 |
ADMA | 2 |
| 2008 | DSCIM: A Novel Service Invocation Mechanism in IMSabstractIn the service invocation mechanism of current IMS (IP multimedia system) network, the serving-call session control function (S-CSCF) invokes each application server (AS) sequentially to perform user's service profile. This mechanism makes SIP request be forwarded excessively among the S-CSCF and application servers (ASs), so it is easy to result in the heavy load of the S-CSCF entity and long call set-up delay. Service capability interaction manager (SCIM) is an entity providing service invoking capability and service interaction management, which is out of standards at present. This paper proposes a novel distributed SCIM (DSCIM) service invocation mechanism within IMS service provision architecture. It aims at reducing the call set-up delay in IMS service layer along with decreasing the load of the S-CSCF entity by invoking each AS consecutively without signaling being forwarded back to the S-CSCF. We model the service invocation mechanisms through Jackson network and the simulations of different scenarios verify that DSCIM service invocation mechanism can effectively reduce call set-up delay of the network and decrease the load level of the S-CSCF. Qi Qi 0001, Jianxin Liao, Xiaomin Zhu 0002, Yufei Cao |
GLOBECOM | 2 |
| 2008 | cmpSCTP: An Extension of SCTP to Support Concurrent Multi-Path TransferabstractThis paper introduced cmpSCTP, a transport layer solution for concurrent multi-path transfer that modifies the standard Stream Control Transmission Protocol (SCTP). The cmpSCTP aims at exploiting SCTP's multi-homing capability by selecting several best paths among multiple available network interfaces to improve data transfer rate to the same multi-homed device. Through the use of path monitoring and packet allotment techniques, cmpSCTP tries to transmit given amount of packets at corresponding path as its ability. At the same time, cmpSCTP updates the transmission strategy based on the real-time information of all of paths. Using cmpSCTP's flexible path management capability, we may switch the flow between multiple paths automatically to realize seamless path handover. Extensive simulations under different scenarios using OPNET verified that cmpSCTP can effectively enhance transmission efficiency and highlighted the superiority of cmpSCTP against the other SCTP's extension implementations under performance indexes such as throughput, handover latency, packet delay, and packet loss. Jianxin Liao, Jingyu Wang 0001, Xiaomin Zhu 0002 |
ICC | 1 |
| 2008 | On Preventing Unnecessary Fast Retransmission With Optimal Fragmentation StrategyabstractMulti-path transferring on the different paths with diverse delays would cause reordering and so unnecessary retransmission which eventually degrades the throughput of connections considerably. To address this issue, we first presented an optimal fragmentation strategy that should be held to avoid unnecessary retransmission/recovery timeouts events in TCP or SCTP due to reordering. Then, we introduced an analytical framework for the optimal fragmentation strategy of real-time traffic. Specifically, this paper formulates optimal traffic fragmentation as a constrained optimization problem using deterministic network calculus and derives its closed-form solution by our proposed progressive water-filling algorithm. Compared with previous work, the proposed scheme is simpler to implement and enforce. Jingyu Wang 0001, Jianxin Liao, Xiaomin Zhu 0002 |
ICC | 2 |
| 2008 | SWPM: An Incremental Fault Localization Algorithm Based on Sliding Window with Preprocessing MechanismabstractMost fault localization techniques are based on time windows. The sizes of time windows impact on the accuracy of fault localization greatly. This paper takes weighted bipartite graph as fault propagation model and proposes a heuristic fault localization approach based on sliding window with preprocessing mechanism (SWPM) to alleviate the shortcomings. First, SWPM defines the concept of symptom extension ratio and partitions observed symptoms into three segments: analyzed segment, analyzing segment, preprocessing segment. Then it determines the most probable fault set through incrementally computing Bayesian suspected degree (BSD) of the three segments and combining their results. Simulations show that the algorithm can reduce the impacts on the accuracy affected by improper window sizes. The algorithm which has a polynomial computational complexity can be applied to large scale communication network. Jianxin Liao, Xiaomin Zhu 0002 |
PDCAT | 2 |
| 2008 | A multi-path mechanism for reliable VoIP transmission over wireless networks
Jianxin Liao, Jingyu Wang 0001, Xiaomin Zhu 0002 |
Comput. Networks | 1 |
| 2008 | Latent Handover: A flow-oriented progressive handover mechanism
Jingyu Wang 0001, Jianxin Liao, Xiaomin Zhu 0002 |
Comput. Commun. | 2 |
| 2007 | Web usability: A user-centered design approach
Xiaomin Zhu 0002, Jianxin Liao |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2007 | The IMS: IP multimedia concepts and services in the mobile domain
Xiaomin Zhu 0002, Jianxin Liao |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2007 | Fuzzy logic and the semantic Web
Xiaomin Zhu 0002, Jianxin Liao |
J. Assoc. Inf. Sci. Technol. | 2 |
| 2007 | Grigoris Antoniou and Frank Van Harmelen, Comments on "A Semantic Web Primer", The MIT Press, Cambridge, Massachusetts, London, England (2004)
Jianxin Liao, Xiaomin Zhu 0002 |
J. Syst. Softw. | 1 |
| 2006 | Optimized Integrated Registration Procedure of Mobile IP and SIP with AAA OperationsabstractIn order to meet the requirements of complicated mobility management in future mobility applications, multi-layers and multi-protocols mobility management schemes should be used. Mobile IP and SIP are protocols used as mobility management protocol in network layer and application layer, respectively. AAA is the foundation of network application deployment. When performing AAA operations in Mobile IP and SIP independently, it is very inefficient. We propose an optimized scheme, named OIRPMSA. In this scheme, through cooperation of Mobile IP and SIP when performing AAA operations, times of round trips are reduced. Analysis indicates that this scheme can reduce signaling cost, time delay and improve system performance significantly. Key words: Mobile IP,SIP,AAA, mobility management. Jianxin Liao, Xiao-Ping Wen, Xiaomin Zhu 0002 |
AINA (1) | 2 |
| 2006 | A UDP-Based State-Sharing Mechanism of SIP Transaction Stateful Proxy
Shi-zhang Luo, Jianxin Liao, Xiaomin Zhu 0002 |
WASA | 2 |
| 2006 | Modeling and Analysis for an Enhanced Three-Tier Dynamic Location Management in 3G
Jiachun Wu, Jianxin Liao, Xiaomin Zhu 0002 |
WASA | 3 |
| 2001 | IIN model: modifications and case study
Xiaomin Zhu 0002, Jianxin Liao |
Comput. Networks | 2 |