Jing Wang 0039

dblp:02/736-39 · DBLP profile ↗
← Back
39ranked-venue papers
0as first author
26since 2021 · last 2026
0000-0002-4627-6307ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 10 since 2021Computer networks · 9 · 1 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 VALU: A Benchmark for Video Anomaly Temporal Localization and Understanding at Multiple Semantic Levels
abstract
Yixiao He, Menghao Zhang, Haifeng Sun, Jing Wang, Kangheng Lin, Jinghan Wang, Chenye Xu, Pengfei Ren, Qi Qi, Jingyu Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yixiao He, Menghao Zhang 0004, Haifeng Sun 0001, Jing Wang 0039, Kangheng Lin, Chenye Xu, Pengfei Ren 0001, Qi Qi 0001, Jingyu Wang 0001
ACL (1)4
2026 RecFlow: Unlocking GPU Efficiency for DLRM Inference via Fine-Grained Parallelism and Incremental Batching
Siheng Pan, Shaolong Li, Minwei Zhang, Shuxi Guo, Haifeng Sun 0001, Qi Qi 0001, Zirui Zhuang, Jianxin Liao, Jing Wang 0039
INFOCOM11
2026 Balancing Flow and Collaboration: Exploring Visual Noise Cancellation in Mixed Reality Workspace
abstract
In open-plan offices, visual noise from surrounding people and objects can negatively impact both concentration and mood. Mixed Reality (MR) offers a promising approach to address this challenge by reshaping the workspace. In this paper, we first conducted a survey with 50 office workers to examine the impact of visual noise, identifying common sources of distraction and potential mitigation strategies. Considering the necessity of face-to-face communication in office environments, we designed adaptive user interfaces to strike a balance between deep focus and seamless in-situ collaboration. We utilized Virtual Reality (VR) and Diminished Reality (DR) methods to eliminate visual noise and leveraged face orientation along with a distance threshold to determine collaborative intentions. We developed a prototype system and conducted a user study for evaluation. The results indicate that our system can create a tranquil workspace to foster concentration and workplace well-being, while maintaining necessary in-situ collaboration. These findings provide valuable insights for designing future MR-integrated office environments.
Xiayang Zhou, Xufeng Jian, Linpei Zhang, Haifeng Sun 0001, Qi Qi 0001, Jing Wang 0039, Jingyu Wang 0001
IUI8
2026 Ubi Grip: Ubiquitous Grip-Based Tangible Object Utilization in Augmented Reality
abstract
Tangible Augmented Reality (AR) enhances user immersion in virtual world by providing haptic feedback through physical proxy objects. However, existing approaches primarily focus on selecting proxy objects based on their global physical properties, neglecting the utilization of local features. Besides, the prevailing strategy of mapping one virtual object to a single dedicated physical proxy creates an inherent switching cost, limiting flexibility and efficiency. Additionally, due to challenges such as real-time performance, generalization and occlusion, the vision-based hand-object tracking remains a difficult task. In this paper, we propose Ubi Grip, an universal hand-object interaction framework for creating grip-based tangible AR applications based on the local graspable feature and a comprehensive hand-object interaction attributes methodology. We employ a lightweight object tracking method to perform tracking, utilizing a hand mask filter and transformation strategy to optimize object pose based on the hand-held properties. Moreover, we design a user-defined workflow for grasping tangible objects, allowing users to switch grips and map interactions. We evaluated our system through comprehensive algorithmic benchmarks and a user study. The benchmarks demonstrate our SOTA performance in object pose estimation and generalization, while the user study validates system usability, providing deeper insights.
Xufeng Jian, Guangtian Liu, Xiayang Zhou, Haifeng Sun 0001, Qi Qi 0001, Pengfei Ren 0001, Shan Jiang 0008, Jing Wang 0039, Jianxin Liao, Jingyu Wang 0001
VR10
2026 DiffNBR: Spatio-Temporal Diffusion with Information Bottleneck for Next-Basket Recommendation
abstract
Next Basket Recommendation (NBR) predicts unordered item sets for users' next purchases, crucial for grocery shopping and online retail scenarios. However, existing NBR methods face two critical challenges: (1) They neglect the threshold-meeting phenomenon, such as common add-to-cart behaviors where users intentionally include seemingly irrelevant items to meet promotional or free shipping price thresholds, which is a widespread phenomenon in reality but has not been a dedicated research focus in the recommendation domain; (2) Even though many works study repetitive and exploratory recommendations, they still lack a theoretically analyzable mechanism to directionally decouple these patterns. This may limit the model's performance. To address these issues, we propose DiffNBR, the first diffusion-model-based framework for NBR. DiffNBR employs two denoising diffusion probabilistic models (DDPMs) to jointly learn users' purchasing behaviors in in spatial and temporal dimensions, modeling the latent compositional strategies the dynamic evolution behind phenomena like threshold-meeting. Moreover, we integrate the information bottleneck theory to enforce directional decoupling between repetition and exploration by explicitly regulating information flows. Specifically, our framework constrains the model's generative representation to focus on learning the exploratory patterns, while the habitual repurchase representation is responsible for the repetitive patterns. Extensive experiments on four real-world datasets show that DiffNBR outperforms the state-of-the-art (SOTA) methods.
Haifeng Sun 0001, Qi Qi 0001, Lejian Zhang, Jing Wang 0039, Jingyu Wang 0001
WSDM6
2025 Unveiling Internal Reasoning Modes in LLMs: A Deep Dive into Latent Reasoning vs. Factual Shortcuts with Attribute Rate Ratio
abstract
Yiran Yang, Haifeng Sun, Jingyu Wang, Qi Qi, Zirui Zhuang, Huazheng Wang, Pengfei Ren, Jing Wang, Jianxin Liao. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Haifeng Sun 0001, Jingyu Wang 0001, Qi Qi 0001, Zirui Zhuang, Huazheng Wang, Pengfei Ren 0001, Jing Wang 0039, Jianxin Liao
EMNLP8
2025 Prior-Aware Dynamic Temporal Modeling Framework for Sequential 3D Hand Pose Estimation
Pengfei Ren 0001, Jingyu Wang 0001, Haifeng Sun 0001, Qi Qi 0001, Menghao Zhang 0004, Lei Zhang 0094, Jing Wang 0039, Jianxin Liao
ICCV8
2025 Masked Self-Supervised Learning and Semantic Noise Separation for Video Anomaly Detection
abstract
Recent progress in video anomaly detection assumes that anomalies cannot be effectively reconstructed because they remain unseen during training. However, we observe that most existing methods excessively rely on appearance features, resulting in the accurate reconstruction of anomalies with subtle short-term appearance variations, which we refer to as appearance confusion. Meanwhile, many approaches fail to exploit sufficient semantic distinction, resulting in motion confusion for anomalies with motion patterns similar to normal ones. In this paper, we propose a masked self-supervised learning-based framework, which effectively addresses the two confusions by exploring context-aware motion patterns and discriminative semantic normality representations. First, we introduce reconstructing multi-pattern masked spatiotemporal information to motivate the model to capture motion patterns that focus on long-term context. Then, we design a semantic noise separation network to address motion confusion, facilitating the construction of semantic normality boundaries through semantic-aware separation. Extensive experiments on the Avenue and ShanghaiTech datasets validate the effectiveness of our proposed method.
Menghao Zhang 0004, Lei Zhang 0094, Qi Qi 0001, Haifeng Sun 0001, Pengfei Ren 0001, Bo He 0003, Jing Wang 0039, Jingyu Wang 0001
ICME8
2025 A³-Net: Calibration-Free Multi-View 3D Hand Reconstruction for Enhanced Musical Instrument Learning
abstract
Precise 3D hand posture is essential for learning musical instruments. Reconstructing highly precise 3D hand gestures enables learners to correct and master proper techniques through 3D simulation and Extended Reality. However, exsiting methods typically rely on precisely calibrated multi-camera systems, which are not easily deployable in everyday environments. In this paper, we focus on calibration-free multi-view 3D hand reconstruction in unconstrained scenarios. Establishing correspondences between multi-view images is particularly challenging without camera extrinsics. To address this, we propose A^3-Net, a multi-level alignment framework that utilizes 3D structural representations with hierarchical geometric and explicit semantic information as alignment proxies, facilitating multi-view feature interaction in both 3D geometric space and 2D visual space. Specifically, we first perfrom global geometric alignment to map multi-view features into a canonical space. Subsequently, we aggregate information into predefined sparse and dense proxies to further integrate cross-view semantics through mutual interaction. Finnaly, we perfrom 2D alignment to align projected 2D visual features with 2D observations. Our method achieves state-of-the-art results in the multi-view 3D hand reconstruction task, demonstrating the effectiveness of our proposed framework.
Geng Chen 0006, Xufeng Jian, Pengfei Ren 0001, Jingyu Wang 0001, Haifeng Sun 0001, Qi Qi 0001, Jing Wang 0039, Jianxin Liao
IJCAI8
2025 Rule Meets Learning: Confidence-Aware Multi-View Fusion for Self-Supervised 3D Hand Pose Estimation
abstract
Self-supervised 3D hand pose estimation methods can leverage labeled synthetic data along with unlabeled real-world data for model training, thereby alleviating the reliance on large-scale annotated datasets. Multi-view information fusion is a key factor in the success of these methods. Rule-based fixed fusion methods are simple, efficient, and generalizable, but they neglect the rich visual information in each view. Neural network-based learnable fusion methods can effectively model both intra- and inter-view semantic context, but they tend to overfit to the domain-specific feature of synthetic data and susceptible to interference of domain gaps. In this paper, we decompose multi-view fusion into two components: a learnable confidence estimation stage and a fixed confidence fusion stage. This design not only enables effective use of multi-view semantic cues but also ensures strong cross-domain generalization. To achieve accurate and robust confidence estimation, our method jointly exploits both multi-view pose consistency and pose-to-data consistency. Experiments on three public datasets demonstrate that our approach significantly outperforms existing state-of-the-art self-supervised 3D hand pose estimation methods.
Pengfei Ren 0001, Jingyu Wang 0001, Haifeng Sun 0001, Qi Qi 0001, Jing Wang 0039, Jianxin Liao
ACM Multimedia5
2025 Evaluating and Mitigating Object Hallucination in Large Vision-Language Models: Can They Still See Removed Objects?
abstract
Yixiao He, Haifeng Sun, Pengfei Ren, Jingyu Wang, Huazheng Wang, Qi Qi, Zirui Zhuang, Jing Wang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Yixiao He, Haifeng Sun 0001, Pengfei Ren 0001, Jingyu Wang 0001, Huazheng Wang, Qi Qi 0001, Zirui Zhuang, Jing Wang 0039
NAACL (Long Papers)8
2025 Unified 2D-3D Discrete Priors for Noise-Robust and Calibration-Free Multiview 3D Human Pose Estimation
abstract
Multi-view 3D human pose estimation (HPE) leverages complementary information across views to improve accuracy and robustness. Traditional methods rely on camera calibration to establish geometric correspondences, which is sensitive to calibration accuracy and lacks flexibility in dynamic settings. Calibration-free approaches address these limitations by learning adaptive view interactions, typically leveraging expressive and flexible continuous representations. However, as the multiview interaction relationship is learned entirely from data without constraint, they are vulnerable to noisy input, which can propagate, amplify and accumulate errors across all views, severely corrupting the final estimated pose. To mitigate this, we propose a novel framework that integrates a noise-resilient discrete prior into the continuous representation-based model. Specifically, we introduce the \textit{UniCodebook}, a unified, compact, robust, and discrete representation complementary to continuous features, allowing the model to benefit from robustness to noise while preserving regression capability. Furthermore, we further propose an attribute-preserving and complementarity-enhancing Discrete-Continuous Spatial Attention (DCSA) mechanism to facilitate interaction between discrete priors and continuous pose features. Extensive experiments on three representative datasets demonstrate that our approach outperforms both calibration-required and calibration-free methods, achieving state-of-the-art performance.
Geng Chen 0006, Pengfei Ren 0001, Xufeng Jian, Haifeng Sun 0001, Menghao Zhang 0004, Qi Qi 0001, Zirui Zhuang, Jing Wang 0039, Jianxin Liao, Jingyu Wang 0001
NeurIPS8
2025 Generalizable Hand-Object Modeling from Monocular RGB Images via 3D Gaussians
abstract
Recent advances in hand-object interaction modeling have employed implicit representations, such as Signed Distance Functions (SDF) and Neural Radiance Fields (NeRF) to reconstruct hands and objects with arbitrary topology and photo-realistic detail. However, these methods often rely on dense 3D surface annotations, or are tailored to short clips constrained in motion trajectories and scene contexts, limiting their generalization to diverse environments and movement patterns. In this work, we present HOGS, an adaptively perceptive 3D Gaussian Splatting (3DGS) framework for generalizable hand-object modeling from unconstrained monocular RGB images. By integrating photometric cues from the visual modality with the physically grounded structure of 3D Gaussians, HOGS disentangles inherent geometry from transient lighting and motion-induced appearance changes. This endows hand-object assets with the ability to generalize to unseen environments and dynamic motion patterns. Experiments on two challenging datasets demonstrate that HOGS outperforms state-of-the-art methods in monocular hand-object reconstruction and photo-realistic rendering.
Pengfei Ren 0001, Qi Qi 0001, Haifeng Sun 0001, Zirui Zhuang, Jing Wang 0039, Jianxin Liao, Jingyu Wang 0001
NeurIPS6
2025 FedMI: Reliable and Privacy-Aware Vertical Federated Learning for Anomaly Detection in Distributed Edge Systems
abstract
Anomaly detection in distributed systems faces critical challenges from feature heterogeneity—where incomplete or divergent feature sets across nodes degrade detection relia-bility—and privacy risks under regulations like GDPR. While federated learning (FL) enables collaborative training without raw data sharing, existing solutions fail to address both challenges simultaneously: traditional FL suffers from performance drops under feature-missing scenarios, and differential privacy techniques introduce utility penalties. This paper proposes FedMI, a vertical federated learning framework that achieves provable privacy preservation and robust anomaly detection in feature-heterogeneous environments. FedMI's key innovations include a novel framework for vertical federated learning in anomaly detection for distributed systems that maintains high detection accuracy, mimicking real-world distributed system conditions, and a mutual information-guided training mechanism that quantifies and minimizes privacy leakage during federated updates. Evaluations on healthcare, financial, and industrial sensor datasets demonstrate FedMI's robustness: it achieves performance comparable to centralized methods in F1-score under data-island scenarios while ensuring compliance with privacy constraints. By unifying privacy quantification and robustness to feature heterogeneity, FedMI advances the development of dependable AI-driven monitoring for distributed systems.
Zirui Zhuang, Qi Qi 0001, Haifeng Sun 0001, Shaoxiong Zhu, Xiaoyuan Fu, Jing Wang 0039
SRDS7
2025 DeepZoning: Re-accelerate CNN Inference with Zoning Graph for Heterogeneous Edge Cluster
abstract
Parallelizing CNN inference on heterogeneous edge clusters with data parallelism has gained popularity as a way to meet real-time requirements without sacrificing model accuracy. However, existing algorithms struggle to find optimal parallel granularity for complex CNNS, the structure of which is a directed acyclic graph (DAG) rather than a chain, and the parallel dimension is inflexible. To distribute the workload of modern CNNs on heterogeneous devices is also proven as NP-hard problem. In this article, we introduce DeepZoning , a versatile and cooperative inference framework that combines both model and data parallelism to accelerate CNN inference. DeepZoning employs two algorithms at different levels: (1) a low-level Adaptive Workload Partition algorithm that uses linear programming and takes spatial and channel dimensions into optimization during the search for feature map distribution on heterogeneous devices, and (2) a high-level Model Partition algorithm that finds the optimal model granularity and organizes complex CNNs into sequential zones to balance communication and computation during execution. Our experimental evaluations show that DeepZoning is effective, achieving up to a 3.02× speed improvement on our experimental prototype compared to state-of-the-art algorithms.
Jingyu Wang 0001, Ruilong Ma, Qi Qi 0001, Zirui Zhuang, Jing Wang 0039, Jianxin Liao, Song Guo 0001
ACM Trans. Archit. Code Optim.6
2023 Region-Aware Dynamic Filtering Network for 3D Hand Reconstruction
abstract
3D hand reconstruction from RGB image has attracted a lot of attention due to its crucial role in human-computer interaction. Nevertheless, it is still challenging to perform 3D hand reconstruction under conditions of hand-object interaction due to severe mutual occlusion. Previous methods usually adopt fixed convolution kernel to extract features. We argue that simply sharing the static filter for all regions is impertinent, given that the occlusion degree varies across different regions, resulting in inconsistent visual representations. To address this issue, we proposed Region-aware Dynamic Filtering Network (RDFNet), which dynamically generates convolution kernels based on the features of different regions, thereby adaptively extracting region-related information. Furthermore, we introduce a dynamic receptive field selection mechanism to determine the most appropriate scale for the convolution kernel. For the severely occluded regions, larger receptive field is needed to capture semantic-related features, while the visible regions are mainly concerned with their own local pattern to accumulate spatial-related features and avoid the interference of irrelevant information. Our proposed RDFNet outperforms state-of-the-art methods by a large margin on several challenging hand-object interaction datasets.
Pengfei Ren 0001, Jingyu Wang 0001, Haifeng Sun 0001, Qi Qi 0001, Jing Wang 0039, Jianxin Liao
ECAI6
2023 How Does Diffusion Influence Pretrained Language Models on Out-of-Distribution Data?
abstract
Transformer-based pretrained language models (PLMs) have achieved great success in modern NLP. An important advantage of PLMs is good out-of-distribution (OOD) robustness. Recently, diffusion models have attracted a lot of work to apply diffusion to PLMs. It remains under-explored how diffusion influences PLMs on OOD data. The core of diffusion models is a forward diffusion process which gradually applies Gaussian noise to inputs, and a reverse denoising process which removes noise. The noised input reconstruction is a fundamental ability of diffusion models. We directly analyze OOD robustness by measuring the reconstruction loss, including testing the abilities to reconstruct OOD data, and to detect OOD samples. Experiments are conducted by analyzing different training parameters and data statistical features on eight datasets. It shows that finetuning PLMs with diffusion degrades the reconstruction ability on OOD data. The comparison also shows that diffusion models can effectively detect OOD samples, achieving state-of-the-art performance in most of the datasets with an absolute accuracy improvement up to 18%. These results indicate that diffusion reduces OOD robustness of PLMs.
Huazheng Wang, Daixuan Cheng, Haifeng Sun 0001, Jingyu Wang 0001, Qi Qi 0001, Jianxin Liao, Jing Wang 0039, Cong Liu 0046
ECAI7
2023 Turn on the Right Track: Weakly Supervised Video Moment Retrieval with Self-Improving Query Reconstruction
abstract
Existing weakly-supervised temporal sentence grounding methods typically regard query reconstruction as the pretext task in place of the absent temporal supervision. However, their approaches suffer from two flaws, i.e. insignificant reconstruction and discrepancy in alignment. Insignificant reconstruction indicates the randomly masked words may not be discriminative enough to distinguish the target event from unrelated events in the video. Discrepancy in alignment indicates the incorrect partial alignment built by query reconstruction task. The flaws undermine the reliability of current reconstruction-based methods. To this end, we propose a novel Self-improving Query ReconstrucTion (SQRT) framework for weakly-supervised temporal sentence grounding. To deal with insignificant reconstruction, we devise a key words mining strategy to determine the important words for language grounding. To attain better moment-query alignment, we introduce inter-sample contrast to tackle the partial alignment built by query reconstruction. The self-improving framework utilizes query reconstruction for language grounding and alleviates the discrepancy in alignment, thus turning on the right track. Experiments on two popular datasets show that SQRT achieves state-of-the-art performance on Charades-STA and comparable performance to the state-of-the-art on ActivityNet Captions.
Haifeng Sun 0001, Jiachang Hao, Jing Wang 0039, Qi Qi 0001, Jingyu Wang 0001, Jianxin Liao
ECAI4
2023 Sample-Adapt Fusion Network for RGB-D Hand Detection in the Wild
abstract
RGB and depth modalities provide complementary information, which can be effectively utilized to improve the performance of hand detection in the wild. Most existing fusion-based methods model the channel-wise or spatial-wise cross-modal correlation to exploit the complementary RGB-D information, in which the modeling operations are shared across all input samples. However, the input images show various modes due to the high diversity of scenes in the wild. This inter-sample variance cannot be effectively perceived by static modeling operations shared across all samples. To address this problem, we propose a Sample-Adapt Fusion Network (SAFNet) with Channel Dynamic Refinement Module (CDRM) and Spatial Dynamic Aggregation Module (SDAM) to adaptively model the channel-wise and spatial-wise cross-modal correlation. Specifically, we propose a Multi-kernel Attention Module (MAM) to individually generate attention maps for each input sample by applying learnable weighting operations to multiple convolutional kernels. Our method outperforms state-of-the-art methods on CUG Hand dataset.
Pengfei Ren 0001, Cong Liu 0046, Jing Wang 0039, Haifeng Sun 0001, Qi Qi 0001, Jingyu Wang 0001
ICASSP5
2023 Robust Video Anomaly Detection Framework via Prior Knowledge and Multi-Path Frame Prediction
abstract
Video anomaly detection aims to automatically detect abnormal objects or behaviors. Most existing methods tackle the problem by minimizing the reconstruction errors stemming from the lack of anomalous data, which leads to poor interpretability and robustness. Focus on the context-dependent nature of anomaly detection, a robust unsupervised Video Anomaly Detection framework based on Knowledge and Frame Prediction is proposed, called VAD-KFP. Prior knowledge which contains the context of anomaly is introduced into the multi-path frame prediction network through multi-layer Graph Convolutional Networks. By integrating the prior knowledge to accurately define anomalies, VAD-KFP is robust to different scenarios and is able to recognize the type of anomaly. An extensive range of experiments have been conducted on three benchmarks, the results of which indicate that our method outperforms strong baselines. Specifically, VAD-KFP obtains an AUROC score of 91.6% for the Avenue dataset.
Menghao Zhang 0004, Jingyu Wang 0001, Jing Wang 0039, Qi Qi 0001, Zirui Zhuang, Haifeng Sun 0001
ICASSP3
2023 Multi-order Matched Neighborhood Consistent Graph Alignment in a Union Vector Space
abstract
In this paper, we study the unsupervised plain graph alignment problem, which aims to find node correspondences across two graphs without any side information. The majority of previous works addressed UPGA based on structural information, which will inevitably lead to subgraph isomorphism issues. That is, unaligned nodes could take similar local structural information. To mitigate this issue, we present the Multi-order Matched Neighborhood Consistent (MMNC) which tries to match nodes by aligning the learned node embeddings with only a small number of pseudo alignment seeds. In particular, we extend matched neighborhood consistency (MNC) to vector space and further develop embedding-based MNC (EMNC). By minimizing the EMNC-based loss function, we can utilize the limited pseudo alignment seeds to approximate the orthogonal transformation matrix between two groups of node embeddings with high efficiency and accuracy. Through extensive experiments on public benchmarks, we show that the proposed methods achieve a good balance between alignment accuracy and speed over multiple datasets compared with existing methods.
Wei Tang 0013, Haifeng Sun 0001, Jingyu Wang 0001, Qi Qi 0001, Jing Wang 0039, Hao Yang 0006, Shimin Tao
SIGIR5
2023 Brief Announcement: Accelerate CNN Inference with Zoning Graph at Dynamic Granularity
abstract
Partitioning a CNN and parallel executing inference with multiple IoT devices have gained popularity as a way to meet real-time requirements without sacrificing model accuracy. However, existing algorithms have struggled to find the optimal model partitioning granularity for complex CNNs. Additionally, executing inference with heterogeneous IoT devices is NP-hard when the structure of the CNN is a directed acyclic graph (DAG) rather than a chain. In this paper, we introduce a versatile and cooperative inference framework that combines both model and data parallelism to accelerate CNN inference. DeepZoning employs two algorithms at different levels: (1) a low-level Adaptive Workload Partition algorithm that uses linear programming and takes spatial and channel dimensions into optimization during the search for feature map distribution on heterogeneous devices, and (2) a high-level Model Partition algorithm that finds the optimal model granularity and organizes complex CNNs into sequential zones to balance communication and computation during execution.
Ruilong Ma, Qi Qi 0001, Jingyu Wang 0001, Zirui Zhuang, Jing Wang 0039
SPAA6
2023 TADL: Fault Localization with Transformer-based Anomaly Detection for Dynamic Microservice Systems
abstract
Due to the complexity of microservice architecture, it is difficult to accomplish efficient microservice anomaly detection and localization tasks and achieve the target of high system reliability. For rapid failure recovery and user satisfaction, it is significant to detect and locate anomalies fast and accurately in microservice systems. In this paper, we propose an anomaly detection and localization model based on Transformer, named TADL (Transformer-based Anomaly Detector and Locator), which models the temporal features and dynamically captures container relationships using Transformer with sandwich structure. TADL uses readily available container performance metrics, making it easy to implement in already-running container clusters. Evaluations are conducted on a sock-shop dataset collected from a real microservice system and a publicly available dataset SMD. Empirical studies on the above two datasets demonstrate that TADL can outperform baseline methods in the performance of anomaly detection, the latency of anomaly detection, and the effect of anomalous container localization, which indicates that TADL is useful in maintaining complex and dynamic microservice systems in the real world.
Yuewei Li, Jingyu Wang 0001, Qi Qi 0001, Jing Wang 0039, Jianxin Liao
SANER5
2023 SA-Fusion: Multimodal Fusion Approach for Web-based Human-Computer Interaction in the Wild
abstract
Web-based AR technology has broadened human-computer interaction scenes from traditional mechanical devices and flat screens to the real world, resulting in unconstrained environmental challenges such as complex backgrounds, extreme illumination, depth range differences, and hand-object interaction. The previous hand detection and 3D hand pose estimation methods are usually based on single modality such as RGB or depth data, which are not available in some scenarios in unconstrained environments due to the differences between the two modalities. To address this problem, we propose a multimodal fusion approach, named Scene-Adapt Fusion (SA-Fusion), which can fully utilize the complementarity of RGB and depth modalities in web-based HCI tasks. SA-Fusion can be applied in existing hand detection and 3D hand pose estimation frameworks to boost their performance, and can be further integrated into the prototyping AR system to construct a web-based interactive AR application for unconstrained environments. To evaluate the proposed multimodal fusion method, we conduct two user studies on CUG Hand and DexYCB dataset, to demonstrate its effectiveness in terms of accurately detecting hand and estimating 3D hand pose in unconstrained environments and hand-object interaction.
Pengfei Ren 0001, Cong Liu 0046, Jing Wang 0039, Haifeng Sun 0001, Qi Qi 0001, Jingyu Wang 0001
WWW5
2023 Identifying Users Across Social Media Networks for Interpretable Fine-Grained Neighborhood Matching by Adaptive GAT
abstract
The primary concern of numerous online social media network (SMN) platforms is how to provide users with effective and personalized web services. To achieve this goal, SMN platforms typically begin by collecting user preferences based on user behaviors (e.g., browsing history, posts) or user profiles. However, the effective information about a specific user on a single SMN platform is limited and monotonous, preventing a comprehensive reflection of the user's preferences. Therefore, recognizing anonymous but identical users across two SMNs to integrate their information is crucial for enhancing web services. Clearly, cross-platform research has the potential to aid in the resolution of numerous problems in service computing theory and applications. Therefore, in this article, we present theCross-PlatformUserMatcher (CPUM) framework, which attempts to map users into a union vector space and then performs user matching based on distance metrics. In particular, we introduce a GNN-based encoderAdaptiveGraphAttention Network (AdaGAT) for modeling user attributes and topology jointly in the social networks to capture two typical alignment principles: topology consistency and attribute consistency. Moreover, we derive AdaGAT from the heuristic of the spectral network alignment technique FINAL, which theoretically guarantees AdaGAT's efficacy. To the best of our knowledge, AdaGAT is the first representation-based alignment model to integrate these two alignment principles synergistically. In addition, two position encoding schemes are introduced to prevent alignment confusion that commonly arises with GNN-based alignment models. Extensive experiments on real-world datasets validate the superiority of the proposed framework.
Wei Tang 0013, Haifeng Sun 0001, Jingyu Wang 0001, Cong Liu 0046, Qi Qi 0001, Jing Wang 0039, Jianxin Liao
IEEE Trans. Serv. Comput.6
2022 Modeling Aspect Correlation for Aspect-based Sentiment Analysis via Recurrent Inverse Learning Guidance
abstract
Aspect-based sentiment analysis (ABSA) aims to distinguish sentiment polarity of every specific aspect in a given sentence. Previous researches have realized the importance of interactive learning with context and aspects. However, these methods are ill-studied to learn complex sentence with multiple aspects due to overlapped polarity feature. And they do not consider the correlation between aspects to distinguish overlapped feature. In order to solve this problem, we propose a new method called Recurrent Inverse Learning Guided Network (RILGNet). Our RILGNet has two points to improve the modeling of aspect correlation and the selecting of aspect feature. First, we use Recurrent Mechanism to improve the joint representation of aspects, which enhances the aspect correlation modeling iteratively. Second, we propose Inverse Learning Guidance to improve the selection of aspect feature by considering aspect correlation, which provides more useful information to determine polarity. Experimental results on SemEval 2014 Datasets demonstrate the effectiveness of RILGNet, and we further prove that RILGNet is state-of-the-art method in multiaspect scenarios.
Longfeng Li, Haifeng Sun 0001, Qi Qi 0001, Jingyu Wang 0001, Jing Wang 0039, Jianxin Liao
COLING5
2019 Continuous Bitrate & Latency Control with Deep Reinforcement Learning for Live Video Streaming
abstract
In this paper, we introduce a continuous bitrate control and latency control model for the Live Video Streaming Challenge. Our model is based on Deep Deterministic Policy Gradient, popular on continuous control tasks. Simultaneously, it can take a fine-grained control through continuous control and does not need to discrete the continuous "latency limit", which is a buffer threshold to minimize end-to-end delay by frame skipping. In all considered live video scenarios, our model can provide a better quality of experience with improvements in average QoE of 3.6% than DQN which discrete the "latency limit". Additionally, challenge results show the effectiveness and applicability of the proposed model, which achieved top performance in 3 different networks that include high, low and oscillating throughput, and ranked the second place in the network with medium throughput.
Ruying Hong, Qiwei Shen 0001, Lei Zhang 0094, Jing Wang 0039
ACM Multimedia4
2019 Deep supervised hashing network with integrated regularisation
abstract
Hashing has been widely deployed to approximate nearest neighbour search for large‐scale multimedia retrieval tasks due to storage and retrieval efficiency. State‐of‐the‐art supervised hashing methods for image retrieval construct deep structures to simultaneously learn image representation and generate good hash codes, and the key step among them is simultaneously learned feature representation and binary hash code. Existing methods use similarity and regularity loss to train deep hashing systems, but these two functions usually work together but not cooperative, which may lead to inadequate performance of the whole system. In this study, a new method for training deep hashing system to learn compact binary codes is presented. The deep supervised hashing network with integrated regularisation (DSHIR) system develop the zero division restriction as a new part of the loss function, which settles the problem of cooperatively guiding the system generate similarity preserving binary codes. DSHIR system also modifies the similarity handling loss to better extract features from image data, which promotes the performance compared to existing end‐to‐end deep hashing systems. Experiments show that DSHIR yields about 10 per cent higher mean average precision on CIFAR‐10 dataset, and also promote on other evaluation indexes compared with state‐of‐the‐art systems.
Jianxin Liao, Baoran Li, Jingyu Wang 0001, Qi Qi 0001, Jing Wang 0039
IET Image Process.6
2018 Actor Model Anomaly Detection Using Kernel Principal Component Analysis
Chunze Wang, Jing Wang 0039, Qiwei Shen 0001
ICONIP (4)2
2016 An Approach to Improve the Cooperation between Heterogeneous SDN Overlays
abstract
The overlay network has been widely developed in recent years. There may be various overlays that co-exist with each other upon the same underlying network. These overlays have heterogeneous performance goals, and they will compete for the physical resources, so that a sub-optimal performance of the overlays may be achieved. Moreover, the heterogeneity of the overlays makes them difficult to coordinate with each other to improve their performance. We introduce the concept of SDN to the deployment of overlay network and propose an approach to make the overlays cooperate with each other. A cooperative solution is proposed for co-existing overlays to improve their performance while leveraging their heterogeneous performance goals. Simulations are performed to evaluate the cooperative solution.
Ziteng Cui, Jianxin Liao, Jingyu Wang 0001, Qi Qi 0001, Jing Wang 0039
LCN5
2016 A dual mode self-adaption handoff for multimedia services in mobile cloud computing environment
Jianxin Liao, Qi Qi 0001, Jing Wang 0039, Jingyu Wang 0001, Yufei Cao
Multim. Tools Appl.3
2016 Game-theoretic model of asymmetrical multipath selection in pervasive computing environment
Jingyu Wang 0001, Jianxin Liao, Tonghong Li, Jing Wang 0039
Pervasive Mob. Comput.4
2015 Cooperative traffic management for co-existing overlays
abstract
The overlay network has been widely deployed by Service Providers to provide services. Since there are multiple SPs built upon the same ISP, their overlays are co-existing and may interfere with each other. The selfishness of overlay may lead to sub-optimal performance and traffic arrangement dilemma for overlays. To optimize the performances of overlays and maximize the benefit of SPs, we proposed a cooperative traffic management framework. Several models are applied to analyze and solve the overlay routing problem, the revenue allocation problem, and the coalition formation problem in the framework. Simulations are performed to evaluate the framework.
Ziteng Cui, Jianxin Liao, Jingyu Wang 0001, Qi Qi 0001, Jing Wang 0039
LCN5
2015 A coalitional game approach on improving interactions in multiple overlay environments
Jianxin Liao, Ziteng Cui, Jingyu Wang 0001, Tonghong Li, Qi Qi 0001, Jing Wang 0039
Comput. Networks6
2015 On the collaborations of multiple selfish overlays using multi-path resources
Jingyu Wang 0001, Jianxin Liao, Tonghong Li, Jing Wang 0039
Peer-to-Peer Netw. Appl.4
2014 Cooperative overlay routing in a multiple overlay environment
abstract
Overlay networks have been widely developed over the past few years. More and more overlays are deployed on the top of the same native network, and share the same physical resources. Competing for these physical resources, co-existing overlays may affect each other adversely. It has been showed that by using selfish overlay routing, co-existing overlays would be likely to converge to a Nash equilibrium which is sub-optimal. However, to achieve the global optimal may also cause the performance degradation of certain overlays, which make it hard to realize. Inspired by the Nash bargaining solution, a cooperative method is proposed for two co-existing overlays to achieve a near Pareto optimal. Simulations are performed to evaluate the proposed approach. The results show that the approach is effective and efficient in the multiple overlay networks environment.
Ziteng Cui, Jianxin Liao, Jingyu Wang 0001, Qi Qi 0001, Jing Wang 0039
ICC5
2014 Introducing collaborations for multi-path selection of multiple selfish overlays
abstract
In complex Internet, different overlay flows are likely to share and compete the same congestible resources. We present a game-theoretic study of the selfish strategic collaboration of multiple overlays when they are allowed to use multipath transfer, which is referred as the multipath selection game. Then we consider the equilibrium in this multipath selection game model where selfish players distribute their overlay traffic. Maximization of the utility functions for each overlay is the criterion of optimality. We adopt the objective of throughput maximization to capture the most typical overlay behaviors, and use the usual TCP as the basis of our analysis. We show analytically the existence and uniqueness of Nash equilibria in these games. Furthermore, we find that the loss of efficiency of Nash equilibria can be arbitrarily large if overlays do not have resource limitations. Our simulations confirm effectiveness and TCP-friendliness of multipath transfer for a range of path number and in the presence of multiple overlay traffic.
Jingyu Wang 0001, Jianxin Liao, Jing Wang 0039, Qi Qi 0001, Tonghong Li
ISCC3
2014 Probe-based end-to-end overload control for networks of SIP servers
Jinzhu Wang, Jianxin Liao, Tonghong Li, Jing Wang 0039, Jingyu Wang 0001, Qi Qi 0001
J. Netw. Comput. Appl.4
2012 A distributed end-to-end overload control mechanism for networks of SIP servers
Jianxin Liao, Jinzhu Wang, Tonghong Li, Jing Wang 0039, Jingyu Wang 0001, Xiaomin Zhu 0002
Comput. Networks4