VLDB 2026 Research / reviewers in the wild / expert
Jiaxing Shen
dblp:193/6060
· DBLP profile ↗
63ranked-venue papers
8as first author
54since 2021 · last 2026
0000-0002-0833-0288ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 24 · 2 first-author · 22 since 2021Databases, data management, data science and information retrieval · 12 · 5 first-author · 8 since 2021Artificial intelligence and machine learning · 11 · 2 first-author · 9 since 2021Systems, architecture and hardware · 11 · 10 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ConnSched: Selective Connection Offloading Framework for Accelerating Stateful NFs with DPU
Fuliang Li, Chengxi Gao, Man Hou, Jiaxing Shen |
INFOCOM | 6 |
| 2026 | One-Sketch: A Unified Framework for Per-Flow Cardinality Measurement with Flexible Bias Control
Kejun Guo, Fuliang Li, Jiaxing Shen, Haorui Wan, Man Hou |
INFOCOM | 3 |
| 2026 | ChatGosen: A Network Configuration Synthesis Approach with Semantic-Computation Decoupling
Chunyuan Liu, Zhaokun Tan, Fuliang Li, Bocheng Liang, Jiaxing Shen, Xingwei Wang 0001 |
IWQoS | 5 |
| 2026 | Duet: Towards Efficient and Accurate Sketch-Based Measurement on DPUs
Fuliang Li, Man Hou, Jiaxing Shen, Xingwei Wang 0001 |
IWQoS | 5 |
| 2026 | The Chatbot Knows It's You: Dialogue Attribution in Unauthenticated Human-LLM Sessions
Haoxuan Kou, Xuefeng Liu 0001, Jiaxing Shen |
WWW | 5 |
| 2026 | Teeth-GS: Gaussian Splatting Diffusion with enamel reflectance prior for single-image tooth crown reconstruction
Yanxing Liang, Yinghui Wang 0001, Jinlong Yang 0002, Tao Yan 0001, Jiaxing Shen |
Medical Image Anal. | 6 |
| 2026 | Adaptive Level-Aware Sketch for Efficient Traffic Measurement in Software Switches
Fuliang Li, Kejun Guo, Yuting Liu 0003, Jiaxing Shen, Xingwei Wang 0001, Jiannong Cao 0001 |
IEEE Trans. Computers | 4 |
| 2026 | A Unified Framework for High-Accuracy and Memory-Efficient Per-Flow Cardinality Measurement
Kejun Guo, Fuliang Li, Haorui Wan, Jiaxing Shen, Xingwei Wang 0001, Jiannong Cao 0001 |
IEEE Trans. Netw. | 5 |
| 2026 | A Unified Configuration Framework for Heterogeneous SketchesabstractNetwork measurement sketches enable efficient traffic monitoring but require careful parameter configuration to balance accuracy and memory efficiency. We presentRA-Sketch, a unified framework for generating memory-optimal sketch configurations that satisfy user-defined error constraints across diverse network measurement tasks. Unlike existing approaches that rely on computationally intensive experimental testing, RA-Sketch introduces: 1)Poisson-distributed collision modelingto construct error predictors for both frequency-independent tasks (membership query, heavy-hitter detection, and super-spreader detection) and frequency-dependent tasks (flow size distribution, frequency estimation, and cardinality estimation), eliminating the need for empirical validation; 2) Ahierarchical search strategycombining power-of-two scaling and binary search, reducing iterations through optimized parameter initialization. RA-Sketch supports 10+ sketch architectures including Bloom Filter, Elastic Sketch, HeavyKeeper, MEC Sketch, MRAC, CM Sketch, CO Sketch, gSkt, rSkt1 among others. Evaluations on real-world network traces demonstrate: 1) up to 6–7 orders-of-magnitude faster configuration than benchmark-based methods; 2) Prediction errors are within 10% for heavy-hitter detection and super-spreader detection in most evaluated settings, while prediction errors for membership query, flow size distribution, frequency estimation, and cardinality estimation are close to zero; 3) Memory utilization approaches theoretical minima. The framework’sgenerality and efficiency enable real-time reconfiguration of sketches under dynamic network conditions. Fuliang Li, Kejun Guo, Yuting Liu 0003, Jiaxing Shen, Xingwei Wang 0001, Jiannong Cao 0001 |
IEEE Trans. Netw. | 4 |
| 2025 | FairTP: A Prolonged Fairness Framework for Traffic PredictionabstractTraffic prediction is pivotal in intelligent transportation systems. Existing works focus mainly on improving overall accuracy, overlooking a crucial problem of whether prediction results will lead to biased decisions by transportation authorities. In practice, the uneven deployment of traffic sensors in different urban areas produces imbalanced data, making the traffic prediction model fail in some urban areas and leading to unfair regional decision-making that eventually severely affects equity and quality of residents’ life. Existing fairness machine learning models struggle to maintain fair traffic prediction over prolonged periods. Although these models might achieve fairness at certain time slots, this static fairness will break down as traffic conditions change. To fill this research gap, we investigate prolonged fair traffic prediction, introducing two novel fairness metrics, i.e., region-based static fairness and sensor-based dynamic fairness, tailored to fairness fluctuations over time and across areas. An innovative prolonged fairness traffic prediction framework, namely FairTP, is then proposed. FairTP achieves prolonged fairness by alternating between “sacrifice” and “benefit” the prediction accuracy of each traffic sensor or area, ensuring that the number of these two actions are balanced over time. Specifically, FairTP incorporates a state identification module to discriminate whether the traffic sensors or areas are in a “sacrifice” or “benefit” state, thereby enabling prolonged fairness-aware traffic predictions. Additionally, we devise a state-guided balanced sampling strategy to select training examples to further enhance prediction fairness by mitigating the performance disparities among areas with uneven sensor distribution over time. Extensive experiments in two real-world datasets show that FairTP significantly improves prediction fairness without causing significant accuracy degradation. Jiangnan Xia, Yu Yang 0012, Jiaxing Shen, Senzhang Wang, Jiannong Cao 0001 |
AAAI | 3 |
| 2025 | SGSimEval: A Comprehensive Multifaceted and Similarity-Enhanced Benchmark for Automatic Survey Generation Systems
Beichen Guo, Yu Yang 0012, Ruosong Yang, Jiaxing Shen |
ADMA (2) | 6 |
| 2025 | Dynamic Ring Signature: Towards Provable Anonymity in Blockchain-Based E-Voting
Shan Jiang 0005, Zhehao Huang, Shichang Xuan, Jiaxing Shen, Huakun Huang, Xiaojie Zhu |
ICA3PP (8) | 4 |
| 2025 | WiCell: Multipath-Enabled Application-Transparent Enhancement for Hybrid Wireless NetworksabstractMobile devices’ multiple network interfaces create the foundation for multipath optimization, yet widespread implementation remains challenging. Existing multipath solutions primarily rely on simulation environments or require hardware/software adaptations, limiting real-world adoption. Meanwhile, 89.8% of mobile users prefer WiFi over cellular networks due to data cost concerns, despite cellular networks often providing better stability. We propose WiCell, an application-transparent multipath enhancement system for mobile networks that addresses these challenges through: (1) a deployment-friendly approach requiring no hardware or software modifications; (2) a tunneling-based architecture with edge proxy servers that dynamically adjusts transmission based on real-time link condition feedback; and (3) two operational modes tailored to different user requirements - WiFi-Primary Adaptive Redundancy Mode for balanced performance and data conservation, and Dual Transmission Mode for maximum connection stability. Our comprehensive evaluation demonstrates WiCell’s effectiveness across 30 common applications. In WiFi-Primary Adaptive Redundancy Mode, WiCell achieves a 74.6% reduction in video stalling rate while consuming only 22.12% of the cellular data compared to full redundancy approaches. In Dual Transmission Mode, WiCell reduces maximum latency by 77.4% compared to WiFi-only transmission, with 61.36% lower average latency than WiFi and 65.84% lower than cellular-only connections. These results validate WiCell as a practical solution for enhancing wireless network performance while maintaining application transparency and addressing user concerns about cellular data consumption. Shuo Jia, Fuliang Li, Jiaxing Shen, Xingwei Wang 0001 |
ICCCN | 4 |
| 2025 | SRv6-ALINT: SRv6-based Efficient In-band Network-Wide Telemetry across LANsabstractAs network size continues to increase and data flows between LANs become more frequent, in-band network telemetry methods across LANs require more balanced path lengths and the privacy of telemetry data. Existing schemes have large variance in telemetry path lengths and do not focus on the riskiness of cross-LAN telemetry. In this paper, we design SRv6-ALINT to direct telemetry paths through SRv6 and upload telemetry data at boundary nodes. In the small-scale network case, SRv6-ALINT solve the path generation strategy to get the theoretical optimal value of the path length variance through a solver. In the case of larger scale networks, SRv6-ALINT propose DFS-stitch, an algorithm that guarantees that the path length variance is as small as possible when the number of paths generated is minimized with no duplicate edges and covering the entire network. For segment list compression in SRv6, SRv6-ALINT propose sequential and binary compression algorithms to compress the obtained paths, reduce the segment list depth and decrease the bandwidth. The evaluation shows that the path length variance generated by DFS-stitch is smaller compared to existing schemes. The compressed segment list length of both segment list compression algorithms is reduced to approximately 50% of the uncompressed length. Kaixiang Yu, Fuliang Li, Jiaxing Shen, Xingwei Wang 0001 |
ICCCN | 4 |
| 2025 | RA-Sketch: A Unified Framework for Rapid and Accurate Sketch ConfigurationsabstractNetwork measurement sketches enable efficient traffic monitoring but require careful parameter configuration to balance accuracy and memory efficiency. We present RA-Sketch, a framework for generating memory-optimal sketch configurations that satisfy user-defined error constraints across diverse network measurement tasks. Unlike existing approaches that rely on computationally intensive experimental testing, RA-Sketch introduces: 1) Poisson-distributed collision modeling to construct error predictors for both frequency-independent tasks (membership query, heavy-hitter detection) and frequency-dependent tasks (frequency/cardinality estimation), eliminating the need for empirical validation; 2) A hierarchical search strategy combining power-of-two scaling and binary search, reducing iterations through optimized parameter initialization. RA-Sketch supports 10+ sketch architectures including Bloom Filter, Elastic Sketch, HeavyGuardian, HeavyKeeper, CM/CO Sketch, gSkt, rSkt1 and so on. Evaluations on real-world network traces demonstrate: 1) 6–7 orders of magnitude faster configuration than benchmark-based methods; 2) Prediction errors ≤10% for heavy-hitter detection, while prediction errors for membership query, and frequency/cardinality estimation are close to zero; 3) Memory utilization approaches theoretical minima. The framework’s generality and efficiency enable real-time reconfiguration of sketches under dynamic network conditions. Kejun Guo, Fuliang Li, Yuting Liu 0003, Jiaxing Shen, Xingwei Wang 0001 |
ICNP | 4 |
| 2025 | MEC-Sketch: Memory-Efficient Per-Flow Cardinality Measurement in High-Speed NetworksabstractPer-Flow cardinality measurement in high-speed networks is essential for network security and traffic analysis applications. Flow cardinality refers to the number of distinct elements within a flow, such as the number of unique destination IPs associated with a given source IP. While extensive research has been conducted on single-flow cardinality estimation, achieving accurate per-flow cardinality measurement with real-time performance and low memory overhead remains challenging in large-scale network environments, particularly given the highly skewed distribution of flow cardinalities where mouse flows with smaller cardinalities dominate, and elephant flows with larger cardinalities are fewer. This paper introduces MEC-Sketch, a memory-efficient cardinality estimation data structure that leverages the inherently skewed distribution of flow cardinalities in network traffic. MEC-Sketch employs a dual-component architecture: a heavy part utilizing a majority vote algorithm for precise super-spreader detection, and a light part implementing compact cardinality estimators for memory-efficient measurement of mouse flows. We address two fundamental technical challenges: (1) adapting the majority vote algorithms to operate with cardinality estimators that lack native support for real-time queries, and (2) implementing an effective mapping strategy between large estimators in the heavy part and small estimators in the light part during elephant-mouse flow separation. Comprehensive evaluations on real-world network traces demonstrate that MEC-Sketch significantly outperforms state-of-the-art solutions in terms of estimation accuracy, memory efficiency, and computational performance for both cardinality estimation and super-spreader detection tasks. Kejun Guo, Fuliang Li, Haorui Wan, Jiaxing Shen, Xingwei Wang 0001 |
ICNP | 5 |
| 2025 | SmartTC: A Real-Time ML-Based Traffic Classification with SmartnicabstractReal-time network traffic classification plays a crucial role in ensuring Quality of Service and network security, and machine learning (ML) based methods achieve high classification accuracy but induce significant computational overhead. While SmartNIC solutions can offload classification tasks thus reducing CPU burdens, they still suffer from various limitations, manifested in limited computing capabilities, insufficiency in dynamic load handling and high latency from heterogeneous computing architectures. To solve these problems, we propose SmartTC, with three key designs: (1) SmartTC employs hardware-software co-design to optimize SmartNIC processing power, (2) SmartTC adopts a trafficaware dynamic batch submission strategy that adjusts submission policies based on real-time network load, and (3) SmartTC proposes parallel pipeline scheduling that ensures efficient task execution while minimizing communication overhead. Finally, we implement SmartTC on the BlueField-3 DPU and conduct extensive experiments for evaluations, and comparison results demonstrate that SmartTC significantly outperforms existing solutions. For example, it reduces average traffic classification time by up to$\mathbf{1 6. 8 \%}$under low loads and$\mathbf{9 0. 9 \%}$under high loads. Besides, SmartTC does not affect Bluefield-3 network services, and saves host CPU usage by at least two cores. Lingxiang Hu, Chenyang Hei, Fuliang Li, Chengxi Gao, Jiaxing Shen, Xingwei Wang 0001 |
IWQoS | 5 |
| 2025 | LA-Sketch: An Adaptive Level-Aware Sketch for Efficient Network Traffic MeasurementabstractNetwork traffic measurement is critical for effective network management. Sketch has been proven to be a promising network traffic measurement solution. Considering the skewed distribution of network traffic, where low-frequency mouse flows dominate and high-frequency elephant flows are fewer, recent sketch-based solutions employ hierarchical designs to enhance memory efficiency and accuracy. However, these solutions inevitably introduce additional challenges, including increased memory access overhead, severe hash collisions between elephant and mouse flows, and limited adaptability to dynamic network environments. In this paper, we propose LA-Sketch, an adaptive level-aware data structure. First, LA-Sketch employs a level-aware classifier to intelligently map each flow to its corresponding level, thereby reducing memory access overhead caused by hierarchical designs and mitigating hash collisions between elephant and mouse flows. Second, we introduce an adaptive counter configuration method that dynamically adjusts the number of counters at each level according to diverse network traffic distributions, which theoretically minimizes overall hash collisions. Finally, to adapt to the continuously changing network traffic characteristics, we propose an adaptive online training method that enables LA-Sketch's classifier to maintain high performance using only sketch query values for training, avoiding the significant overhead of massive traffic data collection. Extensive evaluations on two real-world network traces across five measurement tasks demonstrate that LA-Sketch outperforms state-of-the-art hierarchical sketches. Yuting Liu 0003, Kejun Guo, Fuliang Li, Jiaxing Shen, Xingwei Wang 0001 |
IWQoS | 4 |
| 2025 | EPR-Net: Enhanced patch representation network for point cloud normal estimation
Yinghui Wang 0001, Liangyi Huang, Jinlong Yang 0002, Wei Li 0121, Jiaxing Shen, Xiaojuan Ning |
Comput. Aided Des. | 6 |
| 2025 | SMSMO: Learning to generate multimodal summary for scientific papers
Xinyi Zhong, Zusheng Tan, Shen Gao, Jing Li 0034, Jiaxing Shen, Jing-Yu Ji, Jeff K. T. Tang, Billy Chiu |
Knowl. Based Syst. | 5 |
| 2025 | Distributed Sketch Deployment for Software SwitchesabstractNetwork measurement is critical for various network applications, but scaling measurement techniques to the network-wide level is challenging for existing sketch-based solutions. In software switches, centralized deployment provides low resource usage but suffers from poor load balancing. In contrast, collaborative measurement achieves load balancing through flow distribution across software switches but requires high resource usage. This paper presents a novel distributed deployment framework that overcomes the limitations above. First, our framework is lightweight such that it splits sketches into segments and allocates them across forwarding paths to minimize resource usage and achieve load balancing. This also enables per-packet load balancing by distributing computations across software switches. Second, through a novel collaborative strategy, our framework achieves finer-grained flow distribution and further optimizes load balancing. Third, we further optimize load balancing by eliminating the mutual influence among forwarding paths. We evaluate the proposed framework on various network topologies and different sketches. Results indicate our solution matches the load balancing of collaborative measurement while approaching the low resource usage of centralized deployment. Moreover, it achieves superior performance in per-packet load balancing, which is not considered in previous deployment solutions. Kejun Guo, Fuliang Li, Jiaxing Shen, Xingwei Wang 0001, Jiannong Cao 0001 |
IEEE Trans. Computers | 3 |
| 2025 | Performance Characteristics and Guidelines of Offloading Middleboxes Onto BlueField-2 DPUabstractWith the rapid growth in data center network bandwidth far outpacing improvements in CPU performance, traditional software middleboxes running on servers have become inefficient. The emerging data processing units aim to address this by offloading network functions from the CPU. However, as DPUs are still a new technology, there lacks comprehensive evaluation of their capabilities for accelerating middleboxes. This paper benchmarks and analyzes the performance of offloading middleboxes onto the NVIDIA BlueField-2 DPU. Three key DPU capabilities are explored: flow tables offloading, ARM subsystem packet processing, and connection tracking hardware offload. By applying these to implement representative middleboxes for firewall, packet scheduling, and load balancing, their performance is characterized and compared to conventional CPU-based versions. Results reveal the high throughput of flow tables offloading for stateless firewalls, but limitations as pipeline depth increases. Packet scheduling using ARM cores is shown to currently reduce performance versus CPU-based scheduling. Finally, while connection tracking hardware offload boosts load balancer bandwidth, it also weakens connection creation abilities. Key lessons on efficient middleboxes offloading strategies with DPUs are provided to guide further research and development. Overall, this paper offers useful benchmarking and analysis of emerging DPUs for accelerating middleboxes in modern data centers. Fuliang Li, Jiaxing Shen, Xingwei Wang 0001, Jiannong Cao 0001 |
IEEE Trans. Computers | 3 |
| 2025 | FSA-Hash: Flow-Size-Aware Sketch Hashing for Software SwitchesabstractIn modern data centers and enterprise networks, software switches have become critical components for achieving flexible and efficient network management. Due to resource constraints in software switches, sketches have emerged as a promising approach for network traffic measurement. However, their accuracy is often impacted by hash collisions. Existing hash functions treat all collisions equally, failing to account for the differing impacts of collisions involving elephant flows versus mouse flows. We propose FSA-Hash, a novel flow-size-aware hashing scheme that separates elephant flows from each other and from mouse flows, minimizing the most detrimental collisions. FSA-Hash is designed based on two insights: separating elephant flows from mouse flows avoids overestimating mouse flows, while separating elephant flows from each other enables accurate heavy-hitter detection. We implement FSA-Hash using machine learning models trained on network traffic data (LFSA-Hash), and also design a lightweight online variant (OLFSA-Hash) that learns the hash model solely from sketch queries on the software switch, obviating traffic collection overheads. Evaluations across four sketches and two tasks demonstrate FSA-Hash’s superior accuracy over standard hash functions. Moreover, OLFSA-Hash closely matches LFSA-Hash’s performance, making it an attractive option for adaptively refining the hash model without monitoring traffic. Fuliang Li, Kejun Guo, Yiming Lv, Jiaxing Shen, Yuting Liu 0003, Xingwei Wang 0001, Jiannong Cao 0001 |
IEEE Trans. Computers | 4 |
| 2025 | Fed-OGD: Mitigating Straggler Effects in Federated Learning via Orthogonal Gradient DescentabstractFederated Learning (FL) faces challenges due to straggler clients that impede timely parameter uploads, potentially leading to suboptimal global model performance. Existing approaches using synchronous and asynchronous communication suffer from long waiting times or convergence issues. We propose Fed-OGD, a novel asynchronous FL method addressing the straggler problem through gradient orthogonalization. Our approach innovatively frames the straggler issue using catastrophic forgetting theory, viewing stragglers as instances of the global model “forgetting” to aggregate their parameters. Fed-OGD introduces an Orthogonal Gradient Descent (OGD) technique that caches straggler gradients and orthogonalizes the difference between these and current active client gradients. By projecting active gradients onto straggler orthogonal bases and subtracting the resulting components, we obtain orthogonalized gradients guiding the model towards optimality. We provide theoretical convergence guarantees and demonstrate Fed-OGD’s effectiveness through extensive experiments. Our method achieves state-of-the-art performance across multiple datasets among SOTA FL baselines, with notable improvements in non-IID (non-Independent and identically distributed) scenarios: there are few main categories with many samples while other categories hold few samples in a client. Fed-OGD achieves that 16.66% increase in accuracy on CIFAR-10, and significant gains on CIFAR-100 (5.37%), Tiny-ImageNet (38.51%), and AG_NEWS (16.30%). Wei Li 0121, Zicheng Shen, Xiulong Liu 0001, Chuntao Ding, Jiaxing Shen |
IEEE Trans. Computers | 5 |
| 2025 | Take Your Pick: Enabling Effective Distributed Learning Within Low-Dimensional Feature SpaceabstractPersonalized federated learning (PFL) is a popular distributed learning framework that allows clients to have different models and has many applications where clients' data are in different domains, including autonomous driving, traffic surveillance, and medical diagnosis. The typical model of a client in PFL features a global encoder trained by all clients to extract universal features from the raw data and personalized layers (e.g., a classifier) trained using the client's local data. Nonetheless, due to the differences between the data distributions of different clients (also known as, domain gaps), the universal features produced by the global encoder largely encompass numerous components irrelevant to a certain client's local task. Some recent PFL methods address the above problem by personalizing specific parameters within the encoder. However, these methods encounter substantial challenges attributed to the high dimensionality and nonlinearity of neural network parameter space. In contrast, the feature space exhibits a lower dimensionality, providing greater intuitiveness and interpretability as compared to the parameter space. To this end, we propose a novel PFL framework named FedPick. FedPick achieves PFL within the low-dimensional feature space by adaptively selecting task-relevant features for each client from the features generated by the global encoder based on its local data distribution. It presents a more accessible and interpretable implementation of PFL compared to those methods working in the parameter space. Extensive experimental results on multiple cross-domain datasets show that FedPick can effectively select task-relevant features for each client and improve model performance in cross-domain FL. Guogang Zhu, Xuefeng Liu 0001, Shaojie Tang 0001, Jianwei Niu 0002, Xinghao Wu, Jiaxing Shen, Wanyu Lin |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | Prompting Explicit and Implicit Knowledge for Multi-hop Question Answering Based on Human Reading ProcessabstractPre-trained language models (PLMs) leverage chains-of-thought (CoT) to simulate human reasoning and inference processes, achieving proficient performance in multi-hop QA. However, a gap persists between PLMs’ reasoning abilities and those of humans when tackling complex problems. Psychological studies suggest a vital connection between explicit information in passages and human prior knowledge during reading. Nevertheless, current research has given insufficient attention to linking input passages and PLMs’ pre-training-based knowledge from the perspective of human cognition studies. In this study, we introduce a Prompting Explicit and Implicit knowledge (PEI) framework, which uses prompts to connect explicit and implicit knowledge, aligning with human reading process for multi-hop QA. We consider the input passages as explicit knowledge, employing them to elicit implicit knowledge through unified prompt reasoning. Furthermore, our model incorporates type-specific reasoning via prompts, a form of implicit knowledge. Experimental results show that PEI performs comparably to the state-of-the-art on HotpotQA. Ablation studies confirm the efficacy of our model in bridging and integrating explicit and implicit knowledge. Guangming Huang, Cunjin Luo, Jiaxing Shen |
LREC/COLING | 4 |
| 2024 | STZIP-GNN: A Robust Model for Taxi Demand Prediction in Sparse Urban EnvironmentsabstractAccurate prediction of taxi demand is crucial for optimizing urban transportation systems, improving passenger experiences, and ensuring efficient resource allocation. However, few studies pay close attention to the issue of data sparsity, particularly in high temporal-spatial resolution data, which presents a significant challenge due to the large number of zeros that can affect model prediction performance. Traditional methods predominantly rely on historical order data for predictions, often failing to capture dynamic and contextual information effectively. To address this problem, we propose a spatiotemporal Zero-Inflated Poisson Graph Neural Network (STZIP-GNN) to enhance prediction performance. Our approach leverages the Zero-Inflated Poisson (ZIP) distribution to effectively capture the large number of zeros in sparse data and incorporates additional richer data sources, such as crowdsensing geolocation data, to mitigate the impact of data sparsity on the model. By utilizing the representational power of spatiotemporal graph neural networks, our model fits the parameters of the probability distribution, enhancing prediction performance. Experimental results demonstrate that our model outperforms other baseline models and validates its effectiveness on real-world datasets. Jiaxing Shen |
ICPADS | 2 |
| 2024 | Keep Me Updated: An Empirical Study of Proprietary Vendor Blobs in Android FirmwareabstractDespite extensive security research on various Android components, such as kernel or runtime, little attention has been paid to the proprietary vendor blobs within Android firmware. In this paper, we conduct a large-scale empirical study to understand the update patterns and assess the security implications of vendor blobs. We specifically focus on GPU blobs because they are loaded into every process for displaying graphics user interfaces and can affect the entire system’s security. We examine over 13,000 Android firmware releases between January 2018 and April 2024. Our results reveal that device manufacturers often neglect vendor blob updates. About 82% of firmware releases contain outdated GPU blobs (up to 1,281 days). A significant number of blobs also rely on obsolete LLVM core libraries released more than 15 years ago. To analyze their security implications, we develop a performant fuzzer that requires no physical access to mobile devices. We discover 289 security and behavioral bugs within the blobs. We also present a case study demonstrating how these vulnerabilities can be exploited via WebGL. This work underscores the critical security concerns associated with vulnerable vendor blobs and emphasizes the urgent need for timely updates from device manufacturers. Elliott Wen, Jiaxing Shen, Burkhard Wünsche |
ICPADS | 2 |
| 2024 | Effective Network-Wide Traffic Measurement: A Lightweight Distributed Sketch DeploymentabstractNetwork measurement is critical for various network applications, but scaling measurement techniques to the network-wide level is challenging for existing sketch-based solutions. Centralized sketch deployment provides low resource usage but suffers from poor load balancing. In contrast, collaborative measurement achieves load balancing through flow distribution across switches but requires high resource usage. This paper presents a novel lightweight distributed deployment framework that overcomes the limitations above. First, our framework is lightweight such that it splits sketches into segments and allocates them across forwarding paths to minimize resource usage and achieve load balancing. This also enables per-packet load balancing by distributing computations across switches. Second, our framework is also optimized for load balancing by coordinating between flows and enabling finer-grained flow distribution. We evaluate the proposed framework on various network topologies and different sketch deployments. Results indicate our solution matches the load balancing of collaborative measurement while approaching the low resource usage of centralized deployment. Moreover, it achieves superior performance in per-packet load balancing, which is not considered in previous deployment policies. Our work provides efficient distributed sketch deployment to strike a balance between load balancing and resource usage enabling effective network-wide measurement. Fuliang Li, Kejun Guo, Jiaxing Shen, Xingwei Wang 0001 |
INFOCOM | 3 |
| 2024 | Advancing Sketch-Based Network Measurement: A General, Fine-Grained, Bit-Adaptive Sliding Window FrameworkabstractNetwork measurement plays a critical role in numerous network applications that rely on fundamental flow processing tasks such as frequency estimation, heavy hitter detection, and distribution estimation. Sketch has emerged as an efficient approach for network measurement due to its low overhead. However, most sketch-based solutions target static windows while enabling sliding window-based measurement remains an open challenge. This paper introduces two novel general frameworks applicable to diverse sketch models for sliding window-based network measurement: a traditional sliding window framework and a fine-grained flow-level framework. The traditional framework divides the window into parts and uses centralized flushing to remove expired parts. The flow-level framework tracks timestamps to maintain exact flow characteristics over one period, preventing truncation. To optimize memory usage, a bit-wise adaptive allocation algorithm allows dynamic borrowing of unused counter bits. The frameworks are evaluated on sketches for different flow processing tasks. Results show the frameworks are widely generalizable, reduce error substantially compared to existing approaches, and provide more efficient memory usage. Kejun Guo, Fuliang Li, Jiaxing Shen, Xingwei Wang 0001 |
IWQoS | 3 |
| 2024 | Multi-modal transform-based fusion model for new product sales forecasting
Xiangzhen Li, Jiaxing Shen, Wu Lu |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | Density estimation-based method to determine sample size for random sample partition of big data
Yu-Lin He, Jiaxing Shen, Philippe Fournier-Viger, Joshua Zhexue Huang |
Frontiers Comput. Sci. | 3 |
| 2024 | Resisting TUL attack: balancing data privacy and utility on trajectory via collaborative adversarial learning
Yandi Lun, Hao Miao 0001, Jiaxing Shen, Xiang Wang 0015, Senzhang Wang |
GeoInformatica | 3 |
| 2024 | Privacy-preserving human activity sensing: A surveyabstractWith the prevalence of various sensors and smart devices in people’s daily lives, numerous types of information are being sensed. While using such information provides critical and convenient services, we are gradually exposing every piece of our behavior and activities. Researchers are aware of the privacy risks and have been working on preserving privacy while sensing human activities. This survey reviews existing studies on privacy-preserving human activity sensing. We first introduce the sensors and captured private information related to human activities. We then propose a taxonomy to structure the methods for preserving private information from two aspects: individual and collaborative activity sensing. For each of the two aspects, the methods are classified into three levels: signal, algorithm, and system. Finally, we discuss the open challenges and provide future directions. Yanni Yang 0003, Pengfei Hu 0001, Jiaxing Shen, Haiming Cheng, Zhenlin An, Xiulong Liu 0001 |
High Confid. Comput. | 3 |
| 2024 | ViDSOD-100: A New Dataset and a Baseline Model for RGB-D Video Salient Object Detection
Lei Zhu 0003, Jiaxing Shen, Huazhu Fu, Qing Zhang 0006, Liansheng Wang 0002 |
Int. J. Comput. Vis. | 3 |
| 2024 | NetCR: Knowledge-Graph-Based Recommendation Framework for Manual Network ConfigurationabstractNetwork configuration plays a vital role in quality assurance of network services, requiring considerable effort and time. Automatic network configuration approaches are promising due to their capacity to automatically generate and verify configurations. However, these methods suffer from drawbacks, such as generated configuration content being largely unknown to network operators and inefficient for large-scale networks. Manual configuration is thus still the primary way of managing networks. To facilitate editing processes of manual configuration, a network-wide tool for recommending custom keywords is in urgent need. In this article, we propose a keyword recommendation tool that recommends custom keywords across various network devices. We observe that network devices of the same type and role tend to have a unified template and similar configurations, which enables recommending custom configurations between them. However, the vision entails the following three challenges. First, configurations need to be modeled accurately. Second, a wide variety of network protocols need to be supported. Third, relationships between custom keywords might be implicit and difficult to find. To address the challenges, we first built a configuration knowledge graph that could accurately model configurations, extract latent relationships between keywords, and generate explainable recommendations. Then we applied a recommendation framework to the graph for appropriate keyword recommendations. Lastly, to validate the performance, we conduct recommendations on real configurations over 26 000 times. Experimental results indicate that the overall coverage rate for matching expected configurations reaches 79.396%, and the redundancy rate is less than 20%. Zhenbei Guo, Fuliang Li, Jiaxing Shen, Xingwei Wang 0001 |
IEEE Internet Things J. | 3 |
| 2024 | Human-Intent-Driven Cellular Configuration Generation Using Program SynthesisabstractCellular networks are vital for emerging applications like the Metaverse, which impose demanding quality and quantity requirements. This necessitates frequent reconfiguration of both new and existing base stations to balance network service quality (e.g., ultra-low latency and high bandwidth) and resource consumption. Existing data-driven configuration methods learn from historical data, but have two key limitations. First, they yield only approximate solutions, lacking precision. Second, poor bootstrapping for new base stations with previously unobserved attributes. In this paper, we pioneer intent-driven configuration synthesis by designing an intent language and utilizing satisfiability modulo theory (SMT) for cellular networks to enable exact and precise solutions. We formulate synthesis as an SMT problem, permitting verification of precision. First, we cast configuration generation as a program synthesis problem via novel modeling to bridge the intent-configuration gap. Second, we extend SMT synthesis to scale to large networks. However, vanilla SMT approaches have poor scalability. Hence, we propose an optimization using sampling for constraint verification instead of exhaustive forward solving. We also design a domain-specific optimization to prune the sample space and improve efficiency. Experiments on various network scales demonstrate the effectiveness of our proposed SMT-based cellular network configuration synthesis. Fuliang Li, Chenyang Hei, Jiaxing Shen, Qing Li 0006, Xingwei Wang 0001 |
IEEE J. Sel. Areas Commun. | 3 |
| 2024 | Learning Motion-Guided Multi-Scale Memory Features for Video Shadow DetectionabstractNatural images often contain multiple shadow regions, and existing video shadow detection methods tend to fail in fully identifying all shadow regions, since they mainly learned temporal features at single-scale and single memory. In this work, we develop a novel convolutional neural network (CNN) to learn motion-guided multi-scale memory features to obtain multi-scale temporal information based on multiple network memories for boosting video shadow detection. To do so, our network first constructs three memories (i.e., a global memory, a local memory, and a motion memory) to combine spatial context and object motion for detecting shadows. Based on these three memories, we then devise a multi-scale motion-guided long-short transformer (MMLT) module to learn multi-scale temporal and motion memory features for predicting a shadow detection map of the input video frame. Our MMLT module includes a dense-scale long transformer (DLT), a dense-scale short transformer (DST), and a dense-scale motion transformer (DMT) to read three memories for learning multi-scale transformer features. Our DLT, DST, and DMT consist of a set of memory-read pooling attention (MPA) blocks and densely connect these output features of multiple MPA blocks to learn multi-scale transformer features since the scales of these output features are varied. By doing so, we can more accurately identify multiple shadow regions with different sizes from the input video. Moreover, we devise a self-supervised pretext task to pre-training the feature encoder for enhancing the downstream video shadow detection. Experimental results on three benchmark datasets show that our video shadow detection network quantitatively and qualitatively outperforms 26 state-of-the-art methods. Jiaxing Shen, Xin Yang 0011, Huazhu Fu, Qing Zhang 0006, Ping Li 0016, Bin Sheng 0001, Liansheng Wang 0002, Lei Zhu 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Effective Fault Scenario Identification for Communication Networks via Knowledge-Enhanced Graph Neural NetworksabstractFault Scenario Identification (FSI) is a challenging task that aims to automatically identify the fault types in communication networks from massive alarms to guarantee effective fault recoveries. Existing methods are developed based on rules, which are not accurate enough due to the mismatching issue. In this paper, we propose an effective method named Knowledge-Enhanced Graph Neural Network (KE-GNN), the main idea of which is to integrate the advantages of both the rules and GNN. This work is the first work that employs GNN and rules to tackle the FSI task. Specifically, we encode knowledge using propositional logic and map them into a knowledge space. Then, we elaborately design a teacher-student scheme to minimize the distance between the knowledge embedding and the prediction of GNN, integrating knowledge and enhancing the GNN. To validate the performance of the proposed method, we collected and labeled three real-world 5G fault scenario datasets. Extensive evaluation conducted on these datasets indicates that our method achieves the best performance compared with other representative methods, improving the accuracy by up to 8.10%. Furthermore, the proposed method achieves the best performance against a small dataset setting and can be effectively applied to a new carrier site with a different topology structure. Haihong Zhao, Bo Yang 0002, Jiaxu Cui, Qianli Xing 0002, Jiaxing Shen, Fujin Zhu, Jiannong Cao 0001 |
IEEE Trans. Mob. Comput. | 5 |
| 2024 | Personality-affected Emotion Generation in Dialog SystemsabstractGenerating appropriate emotions for responses is essential for dialogue systems to provide human-like interaction in various application scenarios. Most previous dialogue systems tried to achieve this goal by learning empathetic manners from anonymous conversational data. However, emotional responses generated by those methods may be inconsistent, which will decrease user engagement and service quality. Psychological findings suggest that the emotional expressions of humans are rooted in personality traits. Therefore, we propose a new task, Personality-affected Emotion Generation, to generate emotion based on the personality given to the dialogue system and further investigate a solution through the personality-affected mood transition. Specifically, we first construct a daily dialogue dataset, Personality EmotionLines Dataset ( PELD ), with emotion and personality annotations. Subsequently, we analyze the challenges in this task, i.e., (1) heterogeneously integrating personality and emotional factors and (2) extracting multi-granularity emotional information in the dialogue context. Finally, we propose to model the personality as the transition weight by simulating the mood transition process in the dialogue system and solve the challenges above. We conduct extensive experiments on PELD for evaluation. Results suggest that by adopting our method, the emotion generation performance is improved by 13% in macro-F1 and 5% in weighted-F1 from the BERT-base model. Jiannong Cao 0001, Jiaxing Shen, Ruosong Yang, Shuaiqi Liu 0002, Maosong Sun 0001 |
ACM Trans. Inf. Syst. | 3 |
| 2024 | Distributed Semi-Supervised Learning With Consensus Consistency on Edge DevicesabstractDistributed learning has been increasingly studied in edge computing, enabling edge devices to learn a model collaboratively without exchanging their private data. However, existing approaches assume the private data owned by edge devices are all labeled while the reality is that massive private data are unlabeled and remain to be utilized, which leads to suboptimal performance. To overcome this limitation, we study a new practical problem, Distributed Semi-Supervised Learning (DSSL), to learn models collaboratively with mixed private labeled and unlabeled data on each device. We also propose a novel methodDistMatchthat exploits private unlabeled data by self-training on each device with the help of models from neighboring devices. DistMatch generates pseudo-labels for unlabeled data by properly averaging the predictions of these received models. Furthermore, to avoid self-training with wrong pseudo-labels, DistMatch proposes aconsensus consistencyloss to filter pseudo-labels with high consensus and force the output of the trained model to be consistent with these pseudo-labels. Extensive evaluation results via our self-developed testbed indicate the proposed method outperforms all baselines on commonly used image classification benchmark datasets. Hao-Rui Chen, Lei Yang 0024, Xinglin Zhang 0001, Jiaxing Shen, Jiannong Cao 0001 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2023 | Domain Adaptive Pre-trained Model for Mushroom Image Classification
Zhuo Li 0010, Yu Yang 0012, Jiaxing Shen |
ADMA (4) | 4 |
| 2023 | RemoteGesture: Room-scale Acoustic Gesture Recognition for Multiple UsersabstractAs a promising way, controlling smart devices through gestures offers the benefits of non-contact interaction, efficiency and convenience. Previous researches on acoustic-based gesture recognition have mostly focused on near-field gestures within 1 meter and for a single user only. However, such a nearfield sensing scheme is inadequate to meet the growing demands for multi-person human-computer interaction in far-field spaces. In this paper, we present a novel acoustic-based room-scale gesture recognition system that is capable of recognizing gestures simultaneously performed by multi-user. Our approach achieves far-field sensing by examining the relationship between acoustic signal frame length and sensing range, and overcoming a series of practical challenges incurred by far-field sensing. To simultaneously detect and distinguish gestures of multiple persons, we divide the sensing area into multiple beamforming sub-scanning areas and apply binary search to detect multiple users, which allows for an efficient scanning process and facilitates real-time detection. Finally, we conduct a data augmentation scheme to enlarge the training data and apply a lightweight deep learning framework to classify different gestures. Extensive experiments confirm that our system enables multi-user gesture detection and can recognize nine gestures at a distance up to 7 meters. Mi Tian 0005, Yanwen Wang 0001, Zheng Wang 0054, Junhua Situ, Xiaokang Shi, Jiaxing Shen |
SECON | 8 |
| 2023 | MBA-STNet: Bayes-Enhanced Discriminative Multi-Task Learning for Flow PredictionabstractCrowd flow prediction, which aims to predict the in/out flows of different areas of a city, plays an important role in various applications like intelligent transportation. The challenges of this problem lie in both dynamic mobility patterns of crowds and complex spatial-temporal correlations. Meanwhile, crowd flow is highly correlated to and affected by the Origin-Destination (OD) locations of the flow trajectories, which is largely ignored by existing works. In this paper, we study the novel problem of predicting the crowd flow and flow OD simultaneously, and propose a multi-task bayes-enhanced adversarial spatial temporal network entitled MBA-STNet. MBA-STNet adopts a shared-private framework that contains private spatial-temporal encoders, a shared spatial-temporal encoder, and decoders to learn the task-specific features and shared features. To effectively extract discriminative shared features, an adversarial loss on shared feature extraction is incorporated to reduce information redundancy. A Bayesian Heterogeneous Spatio-temporal Attention Network is designed to learn complex spatio-temporal correlations and alleviate data uncertainty. We also design an attentive temporal queue to capture the complex temporal dependency automatically without domain knowledge. Extensive evaluations are conducted over the bike and taxicab trip datasets in New York. The results demonstrate that the proposed MBA-STNet is superior to state-of-the-art methods. Hao Miao 0001, Jiaxing Shen, Jiannong Cao 0001, Jiangnan Xia, Senzhang Wang |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2022 | DSPBooster: Offloading Unmodified Mobile Applications to DSPs for Power-performance Optimal ExecutionabstractMobile cloud computing offloads intensive code to remote servers to improve execution performance and battery lifetime. Unfortunately, it is prone to data breaches and dependent on network connectivity. In light of these issues, we explore the potential of an under-utilized local computing resource: Digital Signal Processors (DSPs). Programmable DSPs are widely equipped in mobile devices and can conduct mathematical operations at high speed and low power. However, existing mobile applications rarely offload computation to DSPs due to two reasons. Firstly, conventional DSP development requires high proficiency in low-level programming languages. Secondly, DSP application deployment involves many complex steps such as kernel memory allocation and remote procedure calls. In this paper, we introduce DSPBooster, a framework to facilitate application offloading to DSPs for power-performance optimal execution. DSPBooster supports unmodified applications implemented in various high-level programming languages. It transparently deploys suitable application functions to DSPs based on runtime measurement and prediction. Implementing such a system entails many technical challenges thanks to DSPs' unique micro-architecture and inter-processor communication mechanism. In this paper, we provide workable solutions and a thorough system evaluation. We show that DSPBooster can provide up to 11 % performance gain and 3 × power reduction. Elliott Wen, Jiaxing Shen |
COMPSAC | 2 |
| 2022 | Modeling feature interactions for context-aware QoS prediction of IoT services
Zengwei Zheng, Jiaxing Shen, Minyi Guo |
Future Gener. Comput. Syst. | 4 |
| 2022 | Preference-Aware Edge Server Placement in the Internet of ThingsabstractWhile it is well understood that edge computing can significantly facilitate IoT-related applications by deploying edge servers close to IoT devices, it also faces many challenges with numerous IoT devices connected and interacted. One of the most important issues is how to efficiently deploy edge servers under a certain budget with the explosive growth of data scale and user base. Existing studies for edge server placement fail to consider user’s query preferences since individual users may be interested in events in particular regions and are keen to receive up-to-date data streams that originate in regions of interest. In this article, we present a preference-aware edge server placement approach that offers better workload distribution in terms of both minimizing query latency and balancing the load of edge servers. To achieve this, we formulate edge server placement with multiobjective optimization as a${p}$-center problem and design two progressive approaches. We first propose quadratic integer programming (QIP) for small-scale data sets. Since the${p}$-center problem is an NP-hard problem, we thus propose a heuristic algorithm named TAKG (TAbu search with$K$-means and Genetic algorithm) for large-scale data sets. To evaluate the utility of the proposed models, we have conducted a comprehensive evaluation on a large data set that is collected by more than 1900 IoT devices during 30 days. Experimental results indicate our approaches outperform all baselines significantly in terms of both query latency and load balancing. Yihao Lin, Zengwei Zheng, Jiaxing Shen, Minyi Guo |
IEEE Internet Things J. | 5 |
| 2022 | An agnostic and efficient approach to identifying features from execution traces
Chun-Tung Li, Jiannong Cao 0001, Chao Ma 0008, Jiaxing Shen, Ka-Ho Wong |
Knowl. Based Syst. | 4 |
| 2022 | Push the Limit of Acoustic Gesture RecognitionabstractWith the flourish of the smart devices and their applications, controlling devices using gestures has attracted increasing attention for ubiquitous sensing and interaction. Recent works use acoustic signals to track hand movement and recognize gestures. However, they suffer from low robustness due to frequency selective fading, interference and insufficient training data. In this work, we propose RobuCIR, a robust contact-free gesture recognition system that can work under different practical impact factors with high accuracy and robustness. RobuCIR adopts frequency-hopping mechanism to mitigate frequency selective fading and avoid signal interference. To further increase system robustness, we investigate a series of data augmentation techniques based on a small volume of collected data to emulate different practical impact factors. The augmented data is used to effectively train neural network models and cope with various influential factors (e.g., gesture speed, distance to transceiver,etc.). Our experiment results show that RobuCIR can recognize 15 gestures and outperform state-of-the-art works in terms of accuracy and robustness. Yanwen Wang 0001, Jiaxing Shen, Yuanqing Zheng |
IEEE Trans. Mob. Comput. | 2 |
| 2022 | User Profiling Based on Nonlinguistic Audio DataabstractUser profiling refers to inferring people’s attributes of interest ( AoIs ) like gender and occupation, which enables various applications ranging from personalized services to collective analyses. Massive nonlinguistic audio data brings a novel opportunity for user profiling due to the prevalence of studying spontaneous face-to-face communication. Nonlinguistic audio is coarse-grained audio data without linguistic content. It is collected due to privacy concerns in private situations like doctor-patient dialogues. The opportunity facilitates optimized organizational management and personalized healthcare, especially for chronic diseases. In this article, we are the first to build a user profiling system to infer gender and personality based on nonlinguistic audio. Instead of linguistic or acoustic features that are unable to extract, we focus on conversational features that could reflect AoIs. We firstly develop an adaptive voice activity detection algorithm that could address individual differences in voice and false-positive voice activities caused by people nearby. Secondly, we propose a gender-assisted multi-task learning method to combat dynamics in human behavior by integrating gender differences and the correlation of personality traits. According to the experimental evaluation of 100 people in 273 meetings, we achieved 0.759 and 0.652 in F1-score for gender identification and personality recognition, respectively. Jiaxing Shen, Jiannong Cao 0001, Oren Lederman, Shaojie Tang 0001, Alex Pentland |
ACM Trans. Inf. Syst. | 1 |
| 2022 | Automated post scoring: evaluating posts with topics and quoted posts in online forum
Ruosong Yang, Jiannong Cao 0001, Jiaxing Shen |
World Wide Web | 4 |
| 2021 | User Profiling based on Nonlinguistic Audio DataabstractUser profiling refers to inferring people's attributes of interest (AoIs) like gender and occupation, which enables various applications ranging from personalized services to collective analyses. Massive nonlinguistic audio data brings a novel opportunity for user profiling due to the prevalence of studying spontaneous face-to-face communication. In this poster, we are the first to build a user profiling system to infer gender and personality based on nonlinguistic audio. Instead of linguistic or acoustic features which are unable to extract, we focus on conversational features that could reflect AoIs. We firstly develop an adaptive voice activity detection algorithm that could address individual differences in voice and false-positive voice activities caused by people nearby. Secondly, we propose a gender-assisted multi-task learning method to combat dynamics in human behavior by integrating gender differences and the correlation of personality traits. The experimental evaluation of 100 people in 273 meetings indicates the superiority of the proposed method in gender identification and personality recognition respectively. Jiaxing Shen, Oren Lederman, Jiannong Cao 0001, Shaojie Tang 0001, Alex Pentland |
ICDE | 1 |
| 2021 | Repetitive Activity Monitoring from Multivariate Time Series: A Generic and Efficient ApproachabstractRepetitive activities like breathing and walking account for a large fraction of human activities. Monitoring these activities with sensing technology plays a vital role in numerous applications ranging from health monitoring to manufacturing management. Over the last decade, traditional machine learning approaches and recent end-to-end deep learning paradigms have achieved massive successes in human activity recognition. However, these approaches are mostly scenario dependent and computationally expensive. Moreover, real-world repetitive activities may have varying time intervals between each repetition, which invalidate existing sliding window methods. In this paper, we propose STEM, a Scalable Template Extraction Method for scenario independent monitoring of repetitive activities with varying intervals. Instead of using sliding windows, we detect and locate the appearance of repeating patterns based on the Matrix Profile. Distributional features are then extracted from the identified patterns such that domain knowledge can be avoided. The approach is efficient and robust as shown by the evaluation on three public datasets, in which around 95% of the undesired computation were eliminated with up to 4% accuracy improvement. It is also generic as demonstrated by a use case of respiration rate estimation using wireless signals. Chun-Tung Li, Jiaxing Shen, Yanni Yang 0003, Jiannong Cao 0001, Milos Stojmenovic |
MASS | 2 |
| 2021 | BaG: Behavior-Aware Group Detection in Crowded Urban Spaces Using WiFi ProbesabstractGroup detection is gaining popularity as it enables variousXzX applications ranging from marketing to urban planning. Existing methods use received signal strength indicator (RSSI) to detect co-located people as groups. However, this approach might have difficulties in crowded urban spaces since many strangers with similar mobility patterns could be identified as groups. Moreover, RSSI is vulnerable to many factors like the human body attenuation and thus is unreliable in crowded scenarios. In this work, we propose a behavior-aware group detection system (BaG). BaG fuses people’s mobility information and smartphone usage behaviors. We observe that people in a group tend to have similar phone usage patterns. Those patterns could be effectively captured by the proposed feature: number of bursts (NoB). Unlike RSSI, NoB is more resilient to environmental changes as it only cares about receiving packets or not. Besides, both mobility and usage patterns correspond to the same underlying grouping information. We propose a detection method based on collective matrix factorization to reveal the hidden associations by factorizing mobility information and usage patterns simultaneously. Experimental results indicate BaG outperforms baseline approaches by$3.97\% \sim 15.79\%$in F-score. The proposed system could also achieve robust and reliable performance in scenarios with different levels of crowdedness. Jiaxing Shen, Jiannong Cao 0001, Xuefeng Liu 0001 |
IEEE Trans. Mob. Comput. | 1 |
| 2020 | EPARS: Early Prediction of At-Risk Students with Online and Offline Learning Behaviors
Yu Yang 0012, Jiannong Cao 0001, Jiaxing Shen, Hongzhi Yin, Xiaofang Zhou 0001 |
DASFAA (2) | 4 |
| 2020 | Push the Limit of Acoustic Gesture RecognitionabstractWith the flourish of the smart devices and their applications, controlling devices using gestures has attracted increasing attention for ubiquitous sensing and interaction. Recent works use acoustic signals to track hand movement and recognize gestures. However, they suffer from low robustness due to frequency selective fading, interference and insufficient training data. In this work, we propose RobuCIR, a robust contact-free gesture recognition system that can work under different usage scenarios with high accuracy and robustness. RobuCIR adopts frequency-hopping mechanism to mitigate frequency selective fading and avoid signal interference. To further increase system robustness, we investigate a series of data augmentation techniques based on a small volume of collected data to emulate different usage scenarios. The augmented data is used to effectively train neural network models and cope with various influential factors (e.g., gesture speed, distance to transceiver, etc.). Our experiment results show that RobuCIR can recognize 15 gestures and outperform state-of-the-art works in terms of accuracy and robustness. Yanwen Wang 0001, Jiaxing Shen, Yuanqing Zheng |
INFOCOM | 2 |
| 2020 | Dual memory network model for sentiment analysis of review text
Jiaxing Shen, Mingyu Derek Ma, Rong Xiang, Qin Lu 0001, Elvira Perez, Chu-Ren Huang |
Knowl. Based Syst. | 1 |
| 2019 | BaG: Behavior-aware Group Detection in Crowded Urban Spaces using WiFi ProbesabstractGroup detection is gaining popularity as it enables various applications ranging from marketing to urban planning. The group information is an important social context which could facilitate a more comprehensive behavior analysis. An example is for retailers to determine the right incentive for potential customers. Existing methods use received signal strength indicator (RSSI) to detect co-located people as groups. However, this approach might have difficulties in crowded urban spaces since many strangers with similar mobility patterns could be identified as groups. Moreover, RSSI is vulnerable to many factors like the human body attenuation and thus is unreliable in crowded scenarios. In this work, we propose a behavior-aware group detection system (BaG). BaG fuses people's mobility information and smartphone usage behaviors. We observe that people in a group tend to have similar phone usage patterns. Those patterns could be effectively captured by the proposed feature: number of bursts (NoB). Unlike RSSI, NoB is more resilient to environmental changes as it only cares about receiving packets or not. Besides, both mobility and usage patterns correspond to the same underlying grouping information. The latent associations between them cannot be fully utilized in conventional detection methods like graph clustering. We propose a detection method based on collective matrix factorization to reveal the hidden associations by factorizing mobility information and usage patterns simultaneously. Experimental results indicate BaG outperforms baseline approaches by in F-score. The proposed system could also achieve robust and reliable performance in scenarios with different levels of crowdedness. Jiaxing Shen, Jiannong Cao 0001, Xuefeng Liu 0001 |
WWW | 1 |
| 2018 | GINA: Group Gender Identification Using Privacy-Sensitive Audio DataabstractGroup gender is essential in understanding social interaction and group dynamics. With the increasing privacy concerns of studying face-to-face communication in natural settings, many participants are not open to raw audio recording. Existing voice-based gender identification methods rely on acoustic characteristics caused by physiological differences and phonetic differences. However, these methods might become ineffective with privacy-sensitive audio for two main reasons. First, compared to raw audio, privacy-sensitive audio contains significantly fewer acoustic features. Moreover, natural settings generate various uncertainties in the audio data. In this paper, we make the first attempt to identify group gender using privacy-sensitive audio. Instead of extracting acoustic features from privacy-sensitive audio, we focus on conversational features including turn-taking behaviors and interruption patterns. However, conversational behaviors are unstable in gender identification as human behaviors are affected by many factors like emotion and environment. We utilize ensemble feature selection and a two-stage classification to improve the effectiveness and robustness of our approach. Ensemble feature selection could reduce the risk of choosing an unstable subset of features by aggregating the outputs of multiple feature selectors. In the first stage, we infer the gender composition (mixed-gender or same-gender) of a group which is used as an additional input feature for identifying group gender in the second stage. The estimated gender composition significantly improves the performance as it could partially account for the dynamics in conversational behaviors. According to the experimental evaluation of 100 people in 273 meetings, the proposed method outperforms baseline approaches and achieves an F1-score of 0.77 using linear SVM. Jiaxing Shen, Oren Lederman, Jiannong Cao 0001, Florian Berg, Shaojie Tang 0001, Alex Pentland |
ICDM | 1 |
| 2018 | SNOW: Detecting Shopping Groups Using WiFiabstractDetecting shopping groups is gaining popularity as it enables various applications ranging from marketing to advertising. Existing methods exploit WiFi probe requests to detect shopping groups by identifying co-located customers. However, the probe request is prone to suffer from device heterogeneity which might pose a severe data sparseness problem. More importantly, we find that a certain amount of shopping groups would separate sometimes which makes traditional methods unreliable. In this paper, we propose a shopping group detection system using WiFi (SNOW). Instead of collecting probe requests, SNOW utilizes the WiFi data from smartphones associated with the deployed access points (APs). We could thus obtain data from different devices and even ensure a data granularity of seconds using Arping. Besides, we exploit an effective heuristic extracted from two observations of shopping group dynamics to improve the detection performance. First, the probability of group separation differs in diverse areas. Second, the proportion of group participation and individual engagement differs in different activities of the mall. Therefore, APs under which shopping groups appear more frequently and barely separate should contribute more in measuring customer similarity. Lastly, we represent the measured similarity into a matrix format and apply matrix factorization with a sparsity constraint to derive grouping results directly. According to our experiments in a large shopping mall, SNOW improves the detection performance of baseline approaches by 13.2% on average. Jiaxing Shen, Jiannong Cao 0001, Xuefeng Liu 0001, Shaojie Tang 0001 |
IEEE Internet Things J. | 1 |
| 2017 | City-Hunter: Hunting Smartphones in Urban AreasabstractThe security issue of public WiFi is gaining more and more concern. By listening to probe requests, an adversary can obtain the SSID list of the APs to which a smartphone previously connected, and utilizes this information to trick the smartphone into associating to it. However, with the enhancement of security level, most smartphones now do not proactively disclose their SSID lists, making these attacks obsolete. In this paper, we propose City-Hunter, an attacker that can lure nearby smartphones without knowing their SSID information. City-Hunter establishes and maintains an SSID database by integrating both offline and online information. Meanwhile, it smartly chooses some SSIDs to hit a smartphone according to the past record and freshness. We evaluate the performance of City-Hunter in different public places. The results demonstrate that City-Hunter is able to successfully hit 12% ~ 18% smartphones without knowing their SSID information, which is about 4 ~ 8 times improvement compared to the similar attacks like KARMA and MANA. Xuefeng Liu 0001, Jiaqi Wen, Shaojie Tang 0001, Jiannong Cao 0001, Jiaxing Shen |
ICDCS | 5 |
| 2017 | GraphLoc: a graph-based method for indoor subarea localization with zero-configuration
Minyi Guo, Jiaxing Shen, Jiannong Cao 0001 |
Pers. Ubiquitous Comput. | 3 |
| 2017 | DMAD: Data-Driven Measuring of Wi-Fi Access Point Deployment in Urban SpacesabstractWireless networks offer many advantages over wired local area networks such as scalability and mobility. Strategically deployed wireless networks can achieve multiple objectives like traffic offloading, network coverage, and indoor localization. To this end, various mathematical models and optimization algorithms have been proposed to find optimal deployments of access points (APs). However, wireless signals can be blocked by the human body, especially in crowded urban spaces. As a result, the real coverage of an on-site AP deployment may shrink to some degree and lead to unexpected dead spots (areas without wireless coverage). Dead spots are undesirable, since they degrade the user experience in network service continuity, on one hand, and, on the other hand paralyze some applications and services like tracking and monitoring when users are in these areas. Nevertheless, it is nontrivial for existing methods to analyze the impact of human beings on wireless coverage. Site surveys are too time consuming and labor intensive to conduct. It is also infeasible for simulation methods to predict the number of on-site people. In this article, we propose DMAD, a Data-driven Measuring of Wi-Fi Access point Deployment, which not only estimates potential dead spots of an on-site AP deployment but also quantifies their severity, using simple Wi-Fi data collected from the on-site deployment and shop profiles from the Internet. DMAD first classifies static devices and mobile devices with a decision-tree classifier. Then it locates mobile devices to grid-level locations based on shop popularities, wireless signal, and visit duration. Last, DMAD estimates the probability of dead spots for each grid during different time slots and derives their severity considering the probability and the number of potential users. The analysis of Wi-Fi data from static devices indicates that the Pearson Correlation Coefficient of wireless coverage status and the number of on-site people is over 0.7, which confirms that human beings may have a significant impact on wireless coverage. We also conduct extensive experiments in a large shopping mall in Shenzhen. The evaluation results demonstrate that DMAD can find around 70% of dead spots with a precision of over 70%. Jiaxing Shen, Jiannong Cao 0001, Xuefeng Liu 0001, Chisheng Zhang |
ACM Trans. Intell. Syst. Technol. | 1 |