VLDB 2026 Research / reviewers in the wild / expert
Shijian Li
dblp:11/2908
· DBLP profile ↗
88ranked-venue papers
3as first author
51since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 38 · 35 since 2021Human-computer interaction and ubiquitous computing · 20 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 8 since 2021Databases, data management, data science and information retrieval · 13 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 10 since 2021Systems, architecture and hardware · 3 · 3 first-author · 1 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EMOD: A Unified EEG Emotion Representation Framework Leveraging V-A Guided Contrastive LearningabstractEmotion recognition from EEG signals is essential for affective computing and has been widely explored using deep learning. While recent deep learning approaches have achieved strong performance on single EEG emotion datasets, their generalization across datasets remains limited due to the heterogeneity in annotation schemes and data formats. Existing models typically require dataset-specific architectures tailored to input structure and lack semantic alignment across diverse emotion labels. To address these challenges, we propose EMOD: A Unified EEG Emotion Representation Framework Leveraging Valence–Arousal (V–A) Guided Contrastive Learning. EMOD learns transferable and emotion-aware representations from heterogeneous datasets by bridging both semantic and structural gaps. Specifically, we project discrete and continuous emotion labels into a unified V–A space and formulate a soft-weighted supervised contrastive loss that encourages emotionally similar samples to cluster in the latent space. To accommodate variable EEG formats, EMOD employs a flexible backbone comprising a Triple-Domain Encoder followed by a Spatial-Temporal Transformer, enabling robust extraction and integration of temporal, spectral, and spatial features. We pretrain EMOD on 8 public EEG datasets and evaluate its performance on three benchmark datasets. Experimental results show that EMOD achieves the state-of-the-art performance, demonstrating strong adaptability and generalization across diverse EEG-based emotion recognition scenarios. Yuning Chen, Sha Zhao, Shijian Li, Gang Pan 0001 |
AAAI | 3 |
| 2026 | Distillation Dynamics: Towards Understanding Feature-Based Distillation in Vision TransformersabstractWhile feature-based knowledge distillation has proven highly effective for compressing CNNs, these techniques unexpectedly fail when applied to Vision Transformers (ViTs), often performing worse than simple logit-based distillation. We provide the first comprehensive analysis of this phenomenon through a novel analytical framework termed as "distillation dynamics", combining frequency spectrum analysis, information entropy metrics, and activation magnitude tracking. Our investigation reveals that ViTs exhibit a distinctive U-shaped information processing pattern: initial compression followed by expansion. We identify the root cause of negative transfer in feature distillation: a fundamental representational paradigm mismatch between teacher and student models. Through frequency-domain analysis, we show that teacher models employ distributed, high-dimensional encoding strategies in later layers that smaller student models cannot replicate due to limited channel capacity. This mismatch causes late-layer feature alignment to actively harm student performance. Our findings reveal that successful knowledge transfer in ViTs requires moving beyond naive feature mimicry to methods that respect these fundamental representational constraints, providing essential theoretical guidance for designing effective ViTs compression strategies. Hui-yuan Tian, Bonan Xu, Shijian Li |
AAAI | 3 |
| 2026 | EEG Agent: A Unified Framework for Automated EEG Analysis Using Large Language ModelsabstractScalable and generalizable analysis of brain activity is essential for advancing both clinical diagnostics and cognitive research. Electroencephalography (EEG), a non-invasive modality with high temporal resolution, has been widely used for brain states analysis. However, most exiting EEG models are usually tailored for single specific tasks, limiting their utility in realistic scenarios where EEG analysis often involves multi-task and continuous reasoning. In this work, we introduce EEG Agent, a general-purpose framework that leverages large language models (LLMs) to schedule and plan multiple tools to automatically complete EEG-related tasks. EEG Agent is capable of performing the key functions: EEG basic information perception, spatiotemporal EEG exploration, EEG event detection, interaction with users, and EEG report generation. To realize the capabilities, we design a toolbox composed of different tools for EEG preprocessing, feature extraction, event detection, etc. These capabilities were evaluated on public datasets, and our EEG Agent can support flexible and interpretable EEG analysis, highlighting its potential for real-world clinical applications. Sha Zhao, Mingyi Peng, Haiteng Jiang, Shijian Li |
AAAI | 5 |
| 2026 | AgriGPT-Omni: A Unified Speech-Vision-Text Framework for Multilingual Agricultural IntelligenceabstractDespite rapid advances in multimodal large language models, agricultural applications remain constrained by the lack of multilingual speech data, unified multimodal architectures, and comprehensive evaluation benchmarks. To address these challenges, we present AgriGPT-Omni, an agricultural omni-framework that integrates speech, vision, and text in a unified framework.(1) First, we construct a scalable data synthesis and collection pipeline that converts agricultural texts and images into training data, resulting in the largest agricultural speech dataset to date, including 492K synthetic and 1.4K real speech samples across six languages.(2) Second, based on this, we train the first agricultural Omni-model via a three-stage paradigm: textual knowledge injection, progressive multimodal alignment, and GRPO-based reinforcement learning, enabling unified reasoning across languages and modalities.(3) We further propose AgriBench-Omni-2K, the first tri-modal benchmark for agriculture, covering diverse speech–vision–text tasks and multilingual slices, with standardized protocols and reproducible tools. Experiments show that AgriGPT-Omni significantly outperforms general-purpose baselines on multilingual and multimodal reasoning as well as real-world speech understanding. Lanfei Feng, Yunkui Chen, Jianyu Zhang 0001, Nueraili Aierken, Shijian Li |
WWW | 8 |
| 2026 | EEGDiffuser: Label-guided EEG signals synthesis via diffusion model for BCI applications
Jiquan Wang, Sha Zhao, Zhiling Luo, Yangxuan Zhou, Shijian Li, Gang Pan 0001 |
Neurocomputing | 5 |
| 2026 | Cross-subject EEG-based emotion recognition leveraging multi-source domain adaptation with curriculum leaning strategy
Sha Zhao, Yitian Liu, Shijian Li, Gang Pan 0001 |
Neurocomputing | 3 |
| 2026 | Knowledge-Enhanced Multi-Level Session Graph Model for Interactive Recommendation through Deep Reinforcement LearningabstractDeep reinforcement learning (DRL), has shown promise in solving intractable challenges in interactive recommendation systems (IRS). In DRL-based interactive recommendation, state modeling is vital for well-capturing users’ continuous interaction behaviors with recommendation systems. To effectively capture the behavior of users, existing works for state modeling have evolved from sequential-based modeling to session-based modeling. However, existing session-based state modeling works in IRS are still not fully explored with premature session models and insufficient fusion for different session features. As a result, they cannot capture complicated session patterns during interaction, leading to significant information loss. In this article, we propose a Knowledge-enhanced Multi-Level Session Graph (KMSG) model for interactive recommendation to address the above challenge. KMSG models the user’s interactive data into multi-level session graphs and effectively encodes the states via graph neural networks. Specifically, a novel 3-level item transition graph is designed to capture the common session patterns and intra-session item transitions. We further utilize the information from the knowledge graph to enhance the item relations in KMSG. We then design an attention-based graph neural network to propagate the information in KMSG. Extensive experiments on four real-world benchmark datasets demonstrate the superiority of KMSG over state-of-the-art baselines and the rationality of our design in KMSG. Longxiang Shi, Shoujin Wang, Qi Zhang 0020, Kui Su, Shijian Li |
ACM Trans. Knowl. Discov. Data | 8 |
| 2025 | Personalized Sleep Staging Leveraging Source-free Unsupervised Domain AdaptationabstractSleep staging is important for monitoring sleep quality and diagnosing sleep-related disorders. Recently, numerous deep learning-based models have been proposed for automatic sleep staging using polysomnography recordings. Most of them are trained and tested on the same labeled datasets which results in poor generalization to unseen target domains. However, they regard the subjects in the target domains as a whole and overlook the individual discrepancies, which limits the model's generalization ability to new patients (i.e., unseen subjects) and plug-and-play applicability in clinics. To address this, we propose a novel Source-Free Unsupervised Individual Domain Adaptation (SF-UIDA) framework for sleep staging, leveraging sequential cross-view contrasting and pseudo-label based fine-tuning. It is actually a two-step subject-specific adaptation scheme, which enables the source model to effectively adapt to newly appeared unlabeled individual without access to the source data. It meets the practical needs in real-world scenarios, where the personalized customization can be plug-and-play applied to new ones. Our framework is applied to three classic sleep staging models and evaluated on three public sleep datasets, achieving the state-of-the-art performance. Yangxuan Zhou, Sha Zhao, Jiquan Wang, Haiteng Jiang, Shijian Li, Benyan Luo, Gang Pan 0001 |
AAAI | 5 |
| 2025 | Enhancing Recommendation with Reliable Multi-profile Alignment and Collaborative-aware Contrastive LearningabstractRecent studies have explored the integration of Large Language Models (LLMs) into recommender systems to enhance the semantic understanding of users and items. While traditional collaborative filtering approaches primarily rely on interaction histories, LLM-enhanced methods attempt to construct comprehensive profiles by leveraging descriptive metadata and user-generated reviews. The semantic representations of these profiles are then aligned with recommender embeddings to enhance the performance of recommender systems. However, the effectiveness of such approaches heavily depends on the quality of the generated profiles, which face several critical challenges: inaccurate profiles, insufficient information and information gap between semantic representations and recommender embeddings. To tackle these challenges, we propose a novel framework with reliable multi-profile alignment and collaborative-aware contrastive learning. Specifically, we introduce a profile generation method combining Chain-of-Thought(CoT) prompting and self-reflection to address the issue of inaccurate profiles. To alleviate the problem of insufficient information, we introduce an interactive profile construction mechanism that aggregates and summarizes common characteristics from users' and items' neighbors in the user-item graph. To bridge the information gap between semantic representations and recommender embeddings, we propose interactive information fusion(IIF), which aggregates semantic representations from neighbors and employs supervised contrastive learning to guide representation learning. Furthermore, we propose a multi-profile alignment framework that aligns recommender embeddings with both basic profiles and interactive profiles through deduplicated contrastive objectives, facilitating effective semantic-behavioral alignment. Extensive experiments on three public datasets and six base recommenders demonstrate that our method consistently outperforms strong LLM-based baselines, achieving an average improvement of 2.93% in Recall@20 and 2.64% in NDCG@20. Jianyu Zhang 0001, Shijian Li |
CIKM | 3 |
| 2025 | Knowledge Graph Pooling and Unpooling for Concept AbstractionabstractKnowledge graph embedding (KGE) aims to embed entities and relations as vectors in a continuous space and has proven to be effective for KG tasks. Recently, graph neural networks (GNN) based KGEs gain much attention due to their strong capability of encoding complex graph structures. However, most GNN-based KGEs are directly optimized based on the instance triples in KGs, ignoring the latent concepts and hierarchies of the entities. Though some works explicitly inject concepts and hierarchies into models, they are limited to predefined concepts and hierarchies, which are missing in a lot of KGs. Thus in this paper, we propose a novel framework with KG Pooling and unpooling and Contrastive Learning (KGPCL) to abstract and encode the latent concepts for better KG prediction. Specifically, with an input KG, we first construct a U-KG through KG pooling and unpooling. KG pooling abstracts the input graph to a smaller graph as a pooled graph, and KG unpooling recovers the input graph from the pooled graph. Then we model the U-KG with relational KGEs to get the representations of entities and relations for prediction. Finally, we propose the local and global contrastive loss to jointly enhance the representation of entities. Experimental results show that our models outperform the KGE baselines on link prediction task. Juan Li 0010, Wen Zhang 0015, Mingchen Tu, Mingyang Chen 0002, Ningyu Zhang 0001, Shijian Li |
COLING | 7 |
| 2025 | Improving Cross-Task Applicability of Parameter Sharing in Cooperative Multi-Agent Reinforcement LearningabstractParameter sharing is a widely adopted approach in cooperative Multi-Agent Reinforcement Learning (MARL), often achieving strong performance. However, its effectiveness can vary, as the policy’s similarity induced by parameter sharing may hinder performance in certain tasks. In this study, we propose a novel framework, termed Composite Shared Policy (CSP), to enhance the cross-task applicability of parameter sharing. CSP is designed to model multiple diverse policies concurrently, thereby introducing inherent policy diversity without relying on task-specific designs. By increasing the differences among the policies of individual agents, CSP effectively mitigates the policy similarity problem commonly associated with parameter sharing. These characteristics collectively enable CSP to improve the cross-task applicability of parameter sharing. To empirically validate the effectiveness of CSP, we implement it based on QMIX, a classic cooperative MARL method, and conduct experiments across two widely used MARL testbeds. The experimental results demonstrate that CSP significantly enhances the cross-task applicability of parameter sharing. Additionally, we conduct ablation studies to evaluate the contributions of each component within CSP. The results highlight that each component plays a critical role in the overall effectiveness of the framework. The source code is available at https://github.com/Yurui-Li/CSP. Jianyu Zhang 0001, Li Zhang 0045, Shijian Li, Gang Pan 0001 |
ECAI | 4 |
| 2025 | Improving Stability of Parameter Sharing in Cooperative Multi-agent Reinforcement Learning
Li Zhang 0045, Shijian Li, Gang Pan 0001 |
ICANN (1) | 3 |
| 2025 | Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation
Jianyu Zhang 0001, Li Zhang 0045, Shijian Li |
ICANN (2) | 3 |
| 2025 | CBraMod: A Criss-Cross Brain Foundation Model for EEG DecodingabstractElectroencephalography (EEG) is a non-invasive technique to measure and record brain electrical activity, widely used in various BCI and healthcare applications. Early EEG decoding methods rely on supervised learning, limited by specific tasks and datasets, hindering model performance and generalizability. With the success of large language models, there is a growing body of studies focusing on EEG foundation models. However, these studies still leave challenges: Firstly, most of existing EEG foundation models employ full EEG modeling strategy. It models the spatial and temporal dependencies between all EEG patches together, but ignores that the spatial and temporal dependencies are heterogeneous due to the unique structural characteristics of EEG signals. Secondly, existing EEG foundation models have limited generalizability on a wide range of downstream BCI tasks due to varying formats of EEG data, making it challenging to adapt to. To address these challenges, we propose a novel foundation model called CBraMod. Specifically, we devise a criss-cross transformer as the backbone to thoroughly leverage the structural characteristics of EEG signals, which can model spatial and temporal dependencies separately through two parallel attention mechanisms. And we utilize an asymmetric conditional positional encoding scheme which can encode positional information of EEG patches and be easily adapted to the EEG with diverse formats. CBraMod is pre-trained on a very large corpus of EEG through patch-based masked EEG reconstruction. We evaluate CBraMod on up to 10 downstream BCI tasks (12 public datasets). CBraMod achieves the state-of-the-art performance across the wide range of tasks, proving its strong capability and generalizability. The source code is publicly available at https://github.com/wjq-learning/CBraMod. Jiquan Wang, Sha Zhao, Zhiling Luo, Yangxuan Zhou, Haiteng Jiang, Shijian Li, Gang Pan 0001 |
ICLR | 6 |
| 2025 | BrainUICL: An Unsupervised Individual Continual Learning Framework for EEG ApplicationsabstractElectroencephalography (EEG) is a non-invasive brain-computer interface technology used for recording brain electrical activity. It plays an important role in human life and has been widely uesd in real life, including sleep staging, emotion recognition, and motor imagery. However, existing EEG-related models cannot be well applied in practice, especially in clinical settings, where new patients with individual discrepancies appear every day. Such EEG-based model trained on fixed datasets cannot generalize well to the continual flow of numerous unseen subjects in real-world scenarios. This limitation can be addressed through continual learning (CL), wherein the CL model can continuously learn and advance over time. Inspired by CL, we introduce a novel Unsupervised Individual Continual Learning paradigm for handling this issue in practice. We propose the BrainUICL framework, which enables the EEG-based model to continuously adapt to the incoming new subjects. Simultaneously, BrainUICL helps the model absorb new knowledge during each adaptation, thereby advancing its generalization ability for all unseen subjects. The effectiveness of the proposed BrainUICL has been evaluated on three different mainstream EEG tasks. The BrainUICL can effectively balance both the plasticity and stability during CL, achieving better plasticity on new individuals and better stability across all the unseen individuals, which holds significance in a practical setting. Yangxuan Zhou, Sha Zhao, Jiquan Wang, Haiteng Jiang, Shijian Li, Gang Pan 0001 |
ICLR | 5 |
| 2025 | Comparison-Based Beam Search for Constructive NCO Approaches
Hui-yuan Tian, Li Zhang 0045, Runhe Huang, Shijian Li |
ICONIP (1) | 5 |
| 2025 | Boosting Zero-Shot Generalization Performance of Constructive NCO Approaches for Large-Scale TSPabstractLarge-Scale instances pose a significant challenge for neural combinatorial optimization (NCO). Given the NP-Hard nature, it is impractical to directly train neural models on large-scale traveling salesman problem (TSP) instances, and zero-shot generalization becomes a promising way. However, due to distribution shift, existing NCO approaches suffer from significantly deteriorated solution quality on large-scale instances. In this paper, we conduct extensive statistical analysis on LEHD, an advanced constructive NCO approach, revealing: 1) distal nodes exhibit redundancy and negatively impact predictions; 2) greedy inference strategy induces cumulative errors that over-shadows performance. To overcome these limitations, we propose Neural-Guided Beam Search with Simplified Context (NGBS-SC), which consists of a KNN-based sampling method that simplifies context at each construction step, and a beam search strategy that dynamically expands search tree in line with neural predictions. Additionally, we introduce a cropping operation to avoid duplicate search. Using LEHD pre-trained on TSP-100 instances as the backbone model, NGBS-SC achieves impressive performance on both synthetic and real-world datasets. It not only outperforms state-of-the-art NCO approaches, but also surpasses the renowned LKH-3 solver (gap: 0.4125% vs. 0.6348%) on TSP-1K instances. Our work highlights the potential of using NCO approaches through zero-shot generalization, bridging the gap between NCO research and practical applications. Shijian Li |
IJCNN | 3 |
| 2025 | ConceptVQ: Visual Model InterpreterabstractWhile convolution neural networks (CNNs) and vision transformers (ViTs) dominate visual representation learning, the growing model depth causes difficulty for interpretability. Although their internal mechanisms are inspiration from human visual sensation, the entire model fails to emulate human perceptual process, particularly the ability to conceptualize abstract visual elements, discover concept compositionality, and leverage analogical reasoning for real-world interaction. Existing feature visualization methods offer post-hoc insights into model perception but struggle to bridge the gap between statistical patterns and symbolic knowledge. Addressing this, we propose ConceptVQ, a symbolic interpreter that distills pretrained features into visual concepts via a learnable codebook, trained through a von Mises-Fisher Vector Quantized Variational Autoencoder (vMF-VQVAE) framework. By unifying multi-granular features to hyperspherical latent prior, ConceptVQ simulates concept abstraction by bottom-up feature clustering and analogy-making by cross-stage Codebook redistribution, mimicking two important cognitive mechanisms. Experiments demonstrate that our framework outperforms standard VQVAE series in both perceptual reconstruction and semantic fidelity, while its codebook establishes direct, interpretable mappings between pretrained features and discrete visual concepts, revealing hierarchical semantics through locally clustered visual patterns. Moreover, our model works as an architecture-agnostic extension which preserves the pretrained model’s semantic capabilities and maintains high compression rates like VQVAE series, enabling efficient adaptation to downstream symbolic tasks. Our work advances interpretability research by unifying representation learning with cognitively inspired symbolic abstraction, offering a pathway toward human-aligned visual representation. Jianyu Zhang 0001, Li Zhang 0045, Shijian Li |
IJCNN | 3 |
| 2025 | Wearable Music2Emotion : Assessing Emotions Induced by AI-Generated Music through Portable EEG-fNIRS Fusion
Sha Zhao, Song Yi, Yangxuan Zhou, Jiadong Pan, Jiquan Wang, Shijian Li, Shurong Dong, Gang Pan 0001 |
ACM Multimedia | 7 |
| 2025 | SPICED: A Synaptic Homeostasis-Inspired Framework for Unsupervised Continual EEG DecodingabstractHuman brain achieves dynamic stability-plasticity balance through synaptic homeostasis, a self-regulatory mechanism that stabilizes critical memory traces while preserving optimal learning capacities. Inspired by this biological principle, we propose SPICED: a neuromorphic framework that integrates the synaptic homeostasis mechanism for unsupervised continual EEG decoding, particularly addressing practical scenarios where new individuals with inter-individual variability emerge continually. SPICED comprises a novel synaptic network that enables dynamic expansion during continual adaptation through three bio-inspired neural mechanisms: (1) critical memory reactivation, which mimics brain functional specificity, selectively activates task-relevant memories to facilitate adaptation; (2) synaptic consolidation, which strengthens these reactivated critical memory traces and enhances their replay prioritizations for further adaptations and (3) synaptic renormalization, which are periodically triggered to weaken global memory traces to preserve learning capacities. The interplay within synaptic homeostasis dynamically strengthens task-discriminative memory traces and weakens detrimental memories. By integrating these mechanisms with continual learning system, SPICED preferentially replays task-discriminative memory traces that exhibit strong associations with newly emerging individuals, thereby achieving robust adaptations. Meanwhile, SPICED effectively mitigates catastrophic forgetting by suppressing the replay prioritization of detrimental memories during long-term continual learning. Validated on three EEG datasets, SPICED show its effectiveness. More importantly, SPICED bridges biological neural mechanisms and artificial intelligence through synaptic homeostasis, providing insights into the broader applicability of bio-inspired principles. Yangxuan Zhou, Sha Zhao, Jiquan Wang, Haiteng Jiang, Shijian Li, Gang Pan 0001 |
NeurIPS | 5 |
| 2025 | Hierarchical Interaction Summarization and Contrastive Prompting for Explainable Recommendations
Shijian Li |
ECML/PKDD (5) | 3 |
| 2025 | FGDC: A fine-grained divide-and-conquer approach for extending NCO to solve large-scale Traveling Salesman Problem
Li Zhang 0045, Shijian Li, Gang Pan 0001 |
Expert Syst. Appl. | 5 |
| 2025 | M-MDD: A multi-task deep learning framework for major depressive disorder diagnosis using EEG
Yilin Wang 0014, Sha Zhao, Haiteng Jiang, Shijian Li, Gang Pan 0001 |
Neurocomputing | 4 |
| 2025 | EEGMamba: An EEG foundation model with Mamba
Jiquan Wang, Sha Zhao, Zhiling Luo, Yangxuan Zhou, Shijian Li, Gang Pan 0001 |
Neural Networks | 5 |
| 2025 | WemiEnv: An Open-Source Reinforcement Learning Platform for WeChat Mini-GamesabstractThe popularity of mobile games has surged in recent years. Along with mobile games, the emergence of mini-games has recently raised attention. Compared to traditional mobile games, mini-games are more lightweight and platform-independent with low development cost, which has attracted thousands of developers and users. WeChat mini-games platform is one of the most popular platforms with over 100 000 mini-games. The diversity and variety of WeChat mini-games make it an ideal platform for training reinforcement learning (RL) agents. In contrast, most of the existing RL benchmark environments are equipped with predetermined games, which are always limited to several genres and lack the utilization of new and diverse games. To utilize the WeChat mini-games for RL research, in this article, we propose WemiEnv, a lightweight, easy-to-use and open-source platform for RL research towards WeChat mini-games. WemiEnv is built on the WeChat developer tools and allows RL agents to interact with the mini-games. WemiEnv also supports user-customized mini-games, requiring users to implement only a few interface functions within WemiEnv API. We also provide six popular mini-games:Space Fighter,Flip, 2048,Flappy Bird,Timberman, andSnakeas ready-to-use tasks. Experiments were conducted with the OpenAI Spinning Up library for RL baselines on the provided tasks to test the usability of WemiEnv. Longxiang Shi, Qianchen Ding, Jingzhe Hou, Canghong Jin, Ye Tao 0001, Jinling Wei, Shijian Li |
IEEE Trans. Games | 8 |
| 2025 | EvoMoE: Evolutionary Mixture-of-Experts for SSVEP-EEG Classification With User-Independent TrainingabstractThe analysis of EEG data in BCI systems captures unique individual characteristics, presenting diverse patterns that deviate from conventional identical distribution assumptions. Therefore, applying AI models directly to brain data becomes challenging due to the non-identical distribution issue. Meanwhile, as user numbers in BCI systems rise, scalable models are crucial to handle the growing data volume. Moreover, the limited availability of individual data necessitates the use of collective data for training, requiring models with strong generalization capabilities. To address these challenges, we propose Evolutionary Mixture of Experts (EvoMoE), a framework leveraging a set of diverse experts to model data from individuals. Users with similar distributions are grouped together, allowing experts to handle EEG data with different distribution types. The gating network of EvoMoE selects experts that closely match the distribution of the current sample, effectively tackling non-identical distribution issues. When encountering an unrecognized distribution, a new expert is introduced to accommodate the new data pattern, ensuring model adaptability. Evaluations on two 40-category BCI Speller datasets demonstrate significant performance improvements over state-of-the-art methods. On the BETA dataset, our online EvoMoE achieves 13.06% increase in accuracy and a 27.24-point increase in high information transfer rate (ITR) compared to the online UI method. The Bench dataset shows 3.64% increase in accuracy and a 10.42-point increase in ITR. These qualities make it a promising solution for practical BCI implementation, while setting the stage for the development of comprehensive biological big models. Jianyu Zhang 0001, Hui-yuan Tian, Shijian Li, Gang Pan 0001 |
IEEE J. Biomed. Health Informatics | 5 |
| 2024 | Generalizable Sleep Staging via Multi-Level Domain AlignmentabstractAutomatic sleep staging is essential for sleep assessment and disorder diagnosis. Most existing methods depend on one specific dataset and are limited to be generalized to other unseen datasets, for which the training data and testing data are from the same dataset. In this paper, we introduce domain generalization into automatic sleep staging and propose the task of generalizable sleep staging which aims to improve the model generalization ability to unseen datasets. Inspired by existing domain generalization methods, we adopt the feature alignment idea and propose a framework called SleepDG to solve it. Considering both of local salient features and sequential features are important for sleep staging, we propose a Multi-level Feature Alignment combining epoch-level and sequence-level feature alignment to learn domain-invariant feature representations. Specifically, we design an Epoch-level Feature Alignment to align the feature distribution of each single sleep epoch among different domains, and a Sequence-level Feature Alignment to minimize the discrepancy of sequential features among different domains. SleepDG is validated on five public datasets, achieving the state-of-the-art performance. Jiquan Wang, Sha Zhao, Haiteng Jiang, Shijian Li, Gang Pan 0001 |
AAAI | 4 |
| 2024 | A Framework for Image Synthesis Using Supervised Contrastive Learning
Jianyu Zhang 0001, Li Zhang 0045, Shijian Li, Gang Pan 0001 |
ICPR (6) | 4 |
| 2024 | A Novel Multi-Pose Person Re-Identification Method Based on Semantic- and Pose-Guided Feature FusionabstractPerson re-identification (ReID) aims to match the query images with images in the gallery. However, ReID traditionally focuses on outdoor scenes and standing pose, neglecting the complexities of different poses and indoor environments. These neglects correspond to some challenges: postural differences and background noise which interfere with feature learning and matching. To address these issues, we propose a novel multi-pose ReID method, Semantic- and Pose-Guided Feature Fusion (SPGFF), which integrates semantic-guided and pose-guided features. Specially, a semantic-guided module is employed to incorporate global contextual semantic information into the feature representation. This global contextual semantic information refers to the comprehensive understanding of the entire image, including the relationships of different pixels. By incorporating this information, even there are significant variations in pose, the model is enabled to focus on the most pertinent parts of the image. Meanwhile, the pose-guided module uses pose feature to cleanly disentangle bodily semantic components and selectively match corresponding body parts. To the best of our knowledge, this is the first work to introduce the concept of multi-pose ReID, we have provided a benchmark to address the issue of a lack of publicly available datasets. We demonstrate the effectiveness of our approach on both public and proprietary datasets, showcasing its potential to significantly improve person re-identification in previously overlooked scenarios. Yuefeng Ma, Deheng Liu, Zhi-Qi Cheng, Shijian Li |
ICTAI | 4 |
| 2024 | CRTGAN: Controllable Road Network Graphs Generation via Transformer based GANabstractAutomated road network generation is in high demand among various applications. With the surge of generative AI, many methods synthesize road network images or graphs based on deep generative models. However, image-based approaches suffer from modeling topological features, while graph-based approaches struggle to encode spatial information. Furthermore, both of them fail to generate road networks with the desired properties. To alleviate these problems, we propose a novel controllable road network graphs generation pipeline to generate road network graphs with expected properties by focusing on both spatial and topological features. Our generation process is founded on modified transformer-based GAN architectures, which enhances the modeling of global features in road network images, such as topology-consistency. The synthesized road network image from the generator is simultaneously evaluated by the discriminator and the reward network, in aspect of spatial and topological supervision respectively, to control the generation process. Furthermore, we design a differentiable image-to-graph extractor applied in the middle of the pipeline. Extensive experiments demonstrate that our proposed pipeline arise generating road network graphs performance with desired properties, achieving 85.43% in traffic convenience metric. Ruihang Li, Shanding Ye, Shijian Li |
IJCNN | 6 |
| 2024 | Multi-depth branch network for efficient image super-resolution
Hui-yuan Tian, Li Zhang 0045, Shijian Li, Gang Pan 0001 |
Image Vis. Comput. | 3 |
| 2024 | CareSleepNet: A Hybrid Deep Learning Network for Automatic Sleep StagingabstractSleep staging is essential for sleep assessment and plays an important role in disease diagnosis, which refers to the classification of sleep epochs into different sleep stages. Polysomnography (PSG), consisting of many different physiological signals, e.g. electroencephalogram (EEG) and electrooculogram (EOG), is a gold standard for sleep staging. Although existing studies have achieved high performance on automatic sleep staging from PSG, there are still some limitations: 1) they focus on local features but ignore global features within each sleep epoch, and 2) they ignore cross-modality context relationship between EEG and EOG. In this paper, we propose CareSleepNet, a novel hybrid deep learning network for automatic sleep staging from PSG recordings. Specifically, we first design a multi-scale Convolutional-Transformer Epoch Encoder to encode both local salient wave features and global features within each sleep epoch. Then, we devise a Cross-Modality Context Encoder based on co-attention mechanism to model cross-modality context relationship between different modalities. Next, we use a Transformer-based Sequence Encoder to capture the sequential relationship among sleep epochs. Finally, the learned feature representations are fed into an epoch-level classifier to determine the sleep stages. We collected a private sleep dataset, SSND, and use two public datasets, Sleep-EDF-153 and ISRUC to evaluate the performance of CareSleepNet. The experiment results show that our CareSleepNet achieves the state-of-the-art performance on the three datasets. Moreover, we conduct ablation studies and attention visualizations to prove the effectiveness of each module and to analyze the influence of each modality. Jiquan Wang, Sha Zhao, Haiteng Jiang, Yangxuan Zhou, Zhenghe Yu, Shijian Li, Gang Pan 0001 |
IEEE J. Biomed. Health Informatics | 7 |
| 2024 | PU-Detector: A PU Learning-based Framework for Real Money Trading Detection in MMORPGabstractMassive multiplayer online role-playing games (MMORPG) have been becoming one of the most popular and exciting online games. In recent years, a cheating phenomenon called real money trading (RMT) has arisen and damaged the fantasy world in many ways. RMT is the sale of in-game items, currency, or even characters to earn real money, breaking the balance of the game economy ecosystem and damaging the game experience. Therefore, some studies have emerged to address the problem of RMT detection. However, they cannot well handle the label uncertainty problem in practice, where there are only labeled RMT samples (positive samples) and unlabeled samples, which could either be RMT samples or normal transactions (negative samples). Meanwhile, the trading relationship between RMTers is modeled in a simple way, leading to some normal transactions being falsely classified as RMT. In this article, we propose PU-Detector, a novel framework based on PU learning (learning from positive and unlabeled data) for RMT detection, considering the fact that there are only labeled RMT samples and other unlabeled transactions. We first automatically estimate the likelihood of one transaction being RMT by developing an improved PU learning method and proposing an assessment rule. Sequentially, we use the estimated likelihood as edge weight to construct a trading graph to learn trader representation. Then, with the trader representations and basic trading features, we detect RMT samples by the improved PU learning method. PU-Detector is evaluated on a large-scale real world dataset consisting of 33,809,956 transaction logs generated by 43,217 unique players. Compared with other approaches, it achieves the state-of-the-art performance and demonstrates its advantages in detecting underlying RMT samples. Yilin Wang 0014, Sha Zhao, Runze Wu 0001, Yuhong Xu, Jianrong Tao, Tangjie Lv, Shijian Li, Zhipeng Hu, Gang Pan 0001 |
ACM Trans. Knowl. Discov. Data | 8 |
| 2024 | Unsupervised Domain Adaptation With Class-Aware Memory AlignmentabstractUnsupervised domain adaptation (UDA) is to make predictions on unlabeled target domain by learning the knowledge from a label-rich source domain. In practice, existing UDA approaches mainly focus on minimizing the discrepancy between different domains by mini-batch training, where only a few instances are accessible at each iteration. Due to the randomness of sampling, such a batch-level alignment pattern is unstable and may lead to misalignment. To alleviate this risk, we propose class-aware memory alignment (CMA) that models the distributions of the two domains by two auxiliary class-aware memories and performs domain adaptation on these predefined memories. CMA is designed with two distinct characteristics: class-aware memories that create two symmetrical class-aware distributions for different domains and two reliability-based filtering strategies that enhance the reliability of the constructed memory. We further design a unified memory-based loss to jointly improve the transferability and discriminability of features in the memories. State-of-the-art (SOTA) comparisons and careful ablation studies show the effectiveness of our proposed CMA. Hui Wang 0107, Liangli Zheng, Hanbin Zhao, Shijian Li, Xi Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | Loan Fraud Users Detection in Online Lending Leveraging Multiple Data ViewsabstractIn recent years, online lending platforms have been becoming attractive for micro-financing and popular in financial industries. However, such online lending platforms face a high risk of failure due to the lack of expertise on borrowers' creditworthness. Thus, risk forecasting is important to avoid economic loss. Detecting loan fraud users in advance is at the heart of risk forecasting. The purpose of fraud user (borrower) detection is to predict whether one user will fail to make required payments in the future. Detecting fraud users depend on historical loan records. However, a large proportion of users lack such information, especially for new users. In this paper, we attempt to detect loan fraud users from cross domain heterogeneous data views, including user attributes, installed app lists, app installation behaviors, and app-in logs, which compensate for the lack of historical loan records. However, it is difficult to effectively fuse the multiple heterogeneous data views. Moreover, some samples miss one or even more data views, increasing the difficulty in fusion. To address the challenges, we propose a novel end-to-end deep multiview learning approach, which encodes heterogeneous data views into homogeneous ones, generates the missing views based on the learned relationship among all the views, and then fuses all the views together to a comprehensive view for identifying fraud users. Our model is evaluated on a real-world large-scale dataset consisting of 401,978 loan records of 228,117 users from January 1, 2019, to September 30, 2019, achieving the state-of-the-art performance. Sha Zhao, Yongrui Huang, Shijian Li, Gang Pan 0001 |
AAAI | 5 |
| 2023 | Session-based Interactive Recommendation via Deep Reinforcement LearningabstractDeep reinforcement learning (DRL), has shown promise in solving intractable challenges in interactive recommendation systems. In DRL-based interactive recommendation, state modeling is crucial for well-capturing users’ continuous interaction behaviors with shopping systems. A user’s multiple continuous interactions in a given time period (e.g., the time from login to log out) naturally constitute a session. However, existing studies often overlook such valuable session structure and characteristics and instead simply treat them as sequences. As a result, they are not able to capture the complex transitions over users’ interactions within or between sessions, leading to significant information loss. To bridge this significant gap, in this paper, we propose Session-based Interactive Recommendation with Graph Neural Networks (SIR-GNN). SIR-GNN models interaction data as sessions and employs novel graph neural networks to capture rich transition patterns among interactions. Specifically, a novel 3-level transition module is well designed to effectively capture common patterns from all sessions, intra-session transitions, and adjacent-item transitions respectively, followed by an attention-based gated graph neural network to model the state representation for SIR well. Extensive experiments on 3 real-world benchmark datasets demonstrate the superiority of SIR-GNN over state-of-the-art baselines and the rationality of our design in SIR-GNN. Longxiang Shi, Shoujin Wang, Qi Zhang 0020, Shijian Li |
ICDM | 7 |
| 2023 | Pyramid-VAE-GAN: Transferring hierarchical latent variables for image inpaintingabstractSignificant progress has been made in image inpainting methods in recent years. However, they are incapable of producing inpainting results with reasonable structures, rich detail, and sharpness at the same time. In this paper, we propose the Pyramid-VAE-GAN network for image inpainting to address this limitation. Our network is built on a variational autoencoder (VAE) backbone that encodes high-level latent variables to represent complicated high-dimensional prior distributions of images. The prior assists in reconstructing reasonable structures when inpainting. We also adopt a pyramid structure in our model to maintain rich detail in low-level latent variables. To avoid the usual incompatibility of requiring both reasonable structures and rich detail, we propose a novel cross-layer latent variable transfer module. This transfers information about long-range structures contained in high-level latent variables to low-level latent variables representing more detailed information. We further use adversarial training to select the most reasonable results and to improve the sharpness of the images. Extensive experimental results on multiple datasets demonstrate the superiority of our method. Our code is available at https://github.com/thy960112/Pyramid-VAE-GAN . Hui-yuan Tian, Li Zhang 0045, Shijian Li, Gang Pan 0001 |
Comput. Vis. Media | 3 |
| 2023 | RM-FSP: Regret minimization optimizes neural fictitious self-play
Li Zhang 0045, Shijian Li, Xili Chen, Gang Pan 0001 |
Neurocomputing | 3 |
| 2023 | Adaptive cooperative exploration for reinforcement learning from imperfect demonstrations
Fuxian Huang, Naye Ji, Huajian Ni, Shijian Li, Xi Li 0001 |
Pattern Recognit. Lett. | 4 |
| 2023 | Spatio-temporal analysis of urban crime leveraging multisource crowdsensed data
Binbin Zhou 0005, Longbiao Chen, Sha Zhao, Fangxun Zhou, Shijian Li, Gang Pan 0001 |
Pers. Ubiquitous Comput. | 5 |
| 2023 | Unsupervised Domain Adaptation for Crime Risk Prediction Across CitiesabstractCrime risk prediction is crucial for city safety and residents’ life quality. However, without labeled data, it is challenging to predict crime risk in cities. Due to municipal regulations and maintenance costs, it is not trivial for many cities to collect high-quality labeled crime data. In particular, some cities have lots of labeled data while others may have few. It has been possible to develop a crime prediction model for a city without labeled crime data by learning knowledge from a city with abundant data. Nevertheless, the inconsistency of relevant context data between cities exacerbates the difficulty of this prediction task. To this end, this article proposes an effective unsupervised domain adaptation model (UDAC) for crime risk prediction across cities while addressing the contexts’ inconsistency issue. More specifically, we first identify several similar source city grids for each target city grid. Based on these source city grids, we then construct auxiliary contexts for the target city, to make contexts consistent between the two cities. A dense convolutional network with unsupervised domain adaptation is designed to learn high-level representations for accurate crime risk prediction and simultaneously learn domain-invariant features for domain adaptation. The effectiveness of our model is verified through extensive experiments using three real-world datasets. Binbin Zhou 0005, Longbiao Chen, Sha Zhao, Shijian Li, Zengwei Zheng, Gang Pan 0001 |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2022 | T-Detector: A Trajectory based Pre-trained Model for Game Bot Detection in MMORPGsabstractGame bots are programmed to automatically play games and illegally obtain profit, seriously affecting game experience of honest players and breaking the balance of game ecosystem. Therefore, bot detection needs to be addressed urgently, especially for MMORPGs, one of the most rapidly expanding genres of games. There have been some studies for bot detection, but the features they used are dependent on specific games and the methods cannot be generalized to other games. In this paper, we propose a trajectory based pre-trained model for game bot detection from game character trajectories and mouse trajectories, named T-Detector, which is independent to specific games and can be generalized to others. More specifically, we propose a pretrain method of LocationTime2Vec to learn representations of trajectories from huge unlabeled samples, which deeply embed spatial and temporal information hidden in trajectories. Moreover, we extract universal features based on behavioral differences in movement trajectories between human players and bots. We design an Angle Pretrain to extract features of turning angle, and propose an attention pooling module to extract features of moving speed and distance. Such features are not dependent on any specific game, enabling T-Detector to be generalized to many MMORPGs. Evaluated by two large-scale real-world datasets of 143,938 samples from two MMORPGs, T-Detector achieves the state-of-the-art performance in bot detection, and demonstrates powerful generalization ability. Sha Zhao, Junwei Fang, Runze Wu 0001, Jianrong Tao, Shijian Li, Gang Pan 0001 |
ICDE | 6 |
| 2022 | Rapid Earthquake Magnitude Estimation Using Deep LearningabstractEarthquake magnitude estimation is one of the critical parts of earthquake early warning systems. It uses the first few seconds of a waveform recorded by an earthquake detection station, which is required to be rapid and accurate. In this paper, we propose a novel framework to estimate magnitude integrated with deep learning, consisting of feature stage and regression stage. In the feature stage, we extract temporal & spatial features by deep learning methods, and combine them with hand-crafted features embedded expert domain knowledge. Then, each earthquake can be represented by a hybrid feature. Therefore, magnitude estimation can be modeled as a regression problem to solve. Our framework is evaluated on 5,503 earthquake records collected in Sichuan province, China. It is found that, learning the temporal & spatial features by deep neural networks is critical for magnitude estimation. The results demonstrate the state-of-the-art performance, compared with other approaches. Sha Zhao, Yizhi Xu, Zhiling Luo, Jin Dong Song, Shijian Li, Gang Pan 0001 |
IJCNN | 6 |
| 2022 | Answering medical questions in Chinese using automatically mined knowledge and deep neural networks: an end-to-end solutionabstractBACKGROUND: Medical information has rapidly increased on the internet and has become one of the main targets of search engine use. However, medical information on the internet is subject to the problems of quality and accessibility, so ordinary users are unable to obtain answers to their medical questions conveniently. As a solution, researchers build medical question answering (QA) systems. However, research on medical QA in the Chinese language lags behind work on English-based systems. This lag is mainly due to the difficulty of constructing a high-quality knowledge base and the underutilization of medical corpora in the Chinese language. RESULTS: This study developed an end-to-end solution to implement a medical QA system for the Chinese language with low cost and time. First, we created a high-quality medical knowledge graph from hospital data (electronic health/medical records) in a nearly automatic manner that trained a supervised model based on data labeled using bootstrapping techniques. Then, we designed a QA system based on a memory-based neural network and attention mechanism. Finally, we trained the system to generate answers from the knowledge base and a QA corpus on the internet. CONCLUSIONS: Bootstrapping and deep neural network techniques can construct a knowledge graph from electronic health/medical records with satisfactory precision and coverage. Our proposed context bridge mechanisms perform training with a variety of language features. Our QA system can achieve state-of-the-art quality in answering medical questions with constrained topics. As we evaluated, complex Chinese language processing techniques, such as segmentation and parsing, were not necessary for practice and complex architectures were not necessary to build the QA system. Lastly, we created an application using our method for internet QA usage. Li Zhang 0045, Shijian Li, Tianyi Liao, Gang Pan 0001 |
BMC Bioinform. | 3 |
| 2022 | Dynamic road crime risk prediction with urban open data
Binbin Zhou 0005, Longbiao Chen, Fangxun Zhou, Shijian Li, Sha Zhao, Gang Pan 0001 |
Frontiers Comput. Sci. | 4 |
| 2022 | Memory-efficient distribution-guided experience sampling for policy consolidation
Fuxian Huang, Weichao Li 0003, Yining Lin, Naye Ji, Shijian Li, Xi Li 0001 |
Pattern Recognit. Lett. | 5 |
| 2022 | Understanding Smartphone Users From Installed App Lists Using Boolean Matrix FactorizationabstractSmartphones are changing humans' lifestyles. Mobile applications (apps) on smartphones serve as entries for users to access a wide range of services in our daily lives. The apps installed on one's smartphone convey lots of personal information, such as demographics, interests, and needs. This provides a new lens to understand smartphone users. However, it is difficult to compactly characterize a user with his/her installed app list. In this article, a user representation framework is proposed, where we model the underlying relations between apps and users with Boolean matrix factorization (BMF). It builds a compact user subspace by discovering basic components from installed app lists. Each basic component encapsulates a semantic interpretation of a series of special-purpose apps, which is a reflection of user needs and interests. Each user is represented by a linear combination of the semantic basic components. With this user representation framework, we use supervised and unsupervised learning methods to understand users, including mining user attributes, discovering user groups, and labeling semantic tags to users. Extensive experiments were conducted on three data subsets from a large-scale real-world dataset for evaluation, each consisting of installed app lists from over 10 000 users. The results demonstrated the effectiveness of our user representation framework. Sha Zhao, Gang Pan 0001, Jianrong Tao, Zhiling Luo, Shijian Li, Zhaohui Wu 0001 |
IEEE Trans. Cybern. | 5 |
| 2021 | Sync-Switch: Hybrid Parameter Synchronization for Distributed Deep LearningabstractStochastic Gradient Descent (SGD) has become the de facto way to train deep neural networks in distributed clusters. A critical factor in determining the training throughput and model accuracy is the choice of the parameter synchronization protocol. For example, while Bulk Synchronous Parallel (BSP) often achieves better converged accuracy, the corresponding training throughput can be negatively impacted by stragglers. In contrast, Asynchronous Parallel (ASP) can have higher throughput, but its convergence and accuracy can be impacted by stale gradients. To improve the performance of synchronization protocol, recent work often focuses on designing new protocols with a heavy reliance on hard-to-tune hyper-parameters. In this paper, we design a hybrid synchronization approach that exploits the benefits of both BSP and ASP, i.e., reducing training time while simultaneously maintaining the converged accuracy. Based on extensive empirical profiling, we devise a collection of adaptive policies that determine how and when to switch between synchronization protocols. Our policies include both offline ones that target recurring jobs and online ones for handling transient stragglers. We implement the proposed policies in a prototype system, called Sync-Switch, on top of TensorFlow, and evaluate the training performance with popular deep learning models and datasets. Our experiments show that Sync-Switch can achieve ASP level training speedup while maintaining similar converged accuracy when comparing to BSP. Moreover, Sync-Switch's elastic-based policy can adequately mitigate the impact from transient stragglers. Shijian Li, Oren Mangoubi, Lijie Xu, Tian Guo 0001 |
ICDCS | 1 |
| 2021 | A Monte Carlo Neural Fictitious Self-Play approach to approximate Nash Equilibrium in imperfect-information dynamic games
Li Zhang 0045, Wei Wang 0011, Ziliang Han, Shijian Li, Gang Pan 0001 |
Frontiers Comput. Sci. | 5 |
| 2021 | Player Behavior Modeling for Enhancing Role-Playing Game EngagementabstractRole-playing games (RPGs) are one of the most exciting and most rapidly expanding genres of online games. Virtual characters that are not controlled by players, have become an integral part, which helps to advance narratives of RPGs. Believable characters can enhance game engagement and further improve player retention. However, game players easily find that most characters' behaviors are limited and improbable, resulting in a less meaningful game experience. In this work, we propose a framework to model game behaviors to learn behavior patterns of human players. Based on the learned behavior patterns, it generates human-like action sequences that can be used for the design of believable virtual characters in RPGs, so as to enhance game engagement. Specifically, considering the influence of game context in behavior patterns, we integrate game context (players' levels and game classes) with actions together to model behaviors. We propose a long-term memory cell on actions and game context to learn the hidden representations. We also introduce an attention mechanism to measure the contribution of the actions previously performed to the next action. Given only one action, our model can generate action sequences by predicting the succeeding action based on the previously generated actions. The model was evaluated on a real-world data set of over 22 000 players and more than 51 million action logs of an RPG game in 21 days. The results demonstrate the state-of-the-art performance. Sha Zhao, Yizhi Xu, Zhiling Luo, Jianrong Tao, Shijian Li, Changjie Fan, Gang Pan 0001 |
IEEE Trans. Comput. Soc. Syst. | 5 |
| 2021 | HisRect: Features from Historical Visits and Recent Tweet for Co-Location JudgementabstractEnabled by smartphones, social media users are increasingly going mobile. This trend fosters various location based services on social media platforms (e.g., Twitter). Many services like friends notification and community detection benefit from co-location judgement, i.e., to decide whether two Twitter users are co-located in some point-of-interest (POI). This problem is challenging due to the limited information in tweets and the lack of explicit geo-tags in tweets that can be used as labeled data. Our approach to this problem is based on a novel concept of HisRect features extracted from users' historical visits and recent tweets: The former has impacts on where a user visits in general, whereas the latter gives more hints about where a user is currently. In practice, labeled data is scarce. Therefore, we design a semi-supervised learning (SSL) framework that leverages unlabeled data to extract HisRect features. Moreover, we employ an embedding neural network layer to process HisRect features of two users, which decides co-location based on the embedding difference between the two features. Our model is extensively evaluated on two large sets of real Twitter data from more than one million users. The experimental results demonstrate that our HisRect features and SSL framework are highly effective at deciding co-locations. In terms of multiple metrics, our approach clearly outperforms alternative approaches using state-of-the-art techniques. Pengfei Li 0005, Hua Lu 0001, Shijian Li, Gang Pan 0001 |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2020 | PERSEUS: Characterizing Performance and Cost of Multi-Tenant Serving for CNN ModelsabstractDeep learning models are increasingly used for end-user applications, supporting both novel features such as facial recognition, and traditional features, e.g. web search. To accommodate high inference throughput, it is common to host a single pre-trained Convolutional Neural Network (CNN) in dedicated cloud-based servers with hardware accelerators such as Graphics Processing Units (GPUs). However, GPUs can be orders of magnitude more expensive than traditional Central Processing Unit (CPU) servers. These resources could also be under-utilized facing dynamic workloads, which may result in inflated serving costs. One potential way to alleviate this problem is by allowing hosted models to share the underlying resources, which we refer to as multi-tenant inference serving. One of the key challenges is maximizing the resource efficiency for multi-tenant serving given hardware with diverse characteristics, models with unique response time Service Level Agreement (SLA), and dynamic inference workloads. In this paper, we present PERSEUS, a measurement framework that provides the basis for understanding the performance and cost trade-offs of multi-tenant model serving. We implemented PERSEUS in Python atop a popular cloud inference server called Nvidia TensorRT Inference Server. Leveraging PERSEUS, we evaluated the inference throughput and cost for serving various models and demonstrated that multi-tenant model serving led to up to 12% cost reduction. Matthew LeMay, Shijian Li, Tian Guo 0001 |
IC2E | 2 |
| 2020 | Characterizing and Modeling Distributed Training with Transient Cloud GPU ServersabstractCloud GPU servers have become the de facto way for deep learning practitioners to train complex models on large-scale datasets. However, it is challenging to determine the appropriate cluster configuration-e.g., server type and number-for different training workloads while balancing the trade-offs in training time, cost, and model accuracy. Adding to the complexity is the potential to reduce the monetary cost by using cheaper, but revocable, transient GPU servers.In this work, we analyze distributed training performance under diverse cluster configurations using CM-DARE, a cloud-based measurement and training framework. Our empirical datasets include measurements from three GPU types, six geographic regions, twenty convolutional neural networks, and thousands of Google Cloud servers. We also demonstrate the feasibility of predicting training speed and overhead using regression-based models. Finally, we discuss potential use cases of our performance modeling such as detecting and mitigating performance bottlenecks. Shijian Li, Robert J. Walls, Tian Guo 0001 |
ICDCS | 1 |
| 2020 | HisRect: Features from Historical Visits and Recent Tweet for Co-Location JudgementabstractThis study explores the problem of co-location judgement, i.e., to decide whether two Twitter users are co-located at some point-of-interest (POI). We extract novel features, named HisRect, from users' historical visits and recent tweets: The former has impact on where a user visits in general, whereas the latter gives more hints about where a user is currently. To alleviate the issue of data scarcity, a semi-supervised learning (SSL) framework is designed to extract HisRect features. Moreover, we use an embedding neural network layer to decide co-location based on the difference between two users' His-Rect features. Extensive experiments on real Twitter data suggest that our HisRect features and SSL framework are highly effective at deciding co-locations. Pengfei Li 0005, Hua Lu 0001, Shijian Li, Gang Pan 0001 |
ICDE | 4 |
| 2020 | Maximum Entropy Reinforcement Learning with Evolution StrategiesabstractEvolution strategies (ES) have recently raised attention in solving challenging tasks with low computation costs and high scalability. However, it is well-known that evolution strategies reinforcement learning (RL) methods suffer from low stability. Without careful consideration, ES methods are sensitive to local optima and are unstable in learning. Therefore, there is an urgent need for improving the stability of ES methods in solving RL problems. In this paper, we propose a simple yet efficient ES method to stabilize the learning. Specifically, we propose a framework to incorporate the maximum entropy reinforcement learning with evolution strategies and derive an efficient entropy calculation method for linear policies. We further present a practical algorithm called maximum entropy evolution policy search based on the proposed framework, which is efficient and stable for policy search in continuous control. Our algorithm shows high stability across different random seeds and can obtain comparable results in performance against some existing derivative-free RL methods on several of the well-known benchmark MuJoCo robotic control tasks. Longxiang Shi, Shijian Li, Longbing Cao, Long Yang 0004, Gang Pan 0001 |
IJCNN | 2 |
| 2020 | Perception-enhancement based task learning and action scheduling for robotic limb in CPS environment
Shijian Li, Minhao Shi, Runhe Huang, Gang Pan 0001 |
Future Gener. Comput. Syst. | 1 |
| 2020 | Gender Profiling From a Single Snapshot of Apps Installed on a Smartphone: An Empirical StudyabstractThe integration of the fifth generation (5G) networks and artificial intelligence (AI) benefits to create a more holistic and better connected ecosystem for industries. User profiling has become an important issue for industries to improve company profit. In the 5G era, smartphone applications have become an indispensable part in our everyday lives. Users determine what apps to install based on their personal needs, interests, and tastes, which is likely shaped by their genders-the behavioral, cultural, or psychological traits typically associated with their sex. It is possible to profile users' gender based simply on a single snapshot of apps installed on their smartphones. With this inference based on easy to access data, we can make smartphone systems more user-friendly, and provide better personalized products and services. In this article, we explore such possibilities through an empirical study on a large-scale dataset of installed app lists from 15 000 Android users. More specifically, we investigate the following research questions: 1) What differences between females and males can be explored from installed app lists? 2) Can user gender be reliably inferred from a snapshot of apps installed? Which snapshot feature(s) are the most predictive? What is the best combination of features for building the gender prediction model? 3) What are the limitations of a gender prediction model based solely on a snapshot of apps installed on a smartphone? We find significant gender differences in app type, function, and icon design. We then extract the corresponding features from a snapshot of apps installed to infer the gender of each user. We assess the gender predictive ability of individual features and combinations of different features. We achieve an accuracy of 76.62% and area under the curve of 84.23% with the best set of features, outperforming the existing work by around 5% and 10%, respectively. Finally, we perform an error analysis on misclassified users and discussed the implications and limitations of this article. Sha Zhao, Yizhi Xu, Xiaojuan Ma, Ziwen Jiang, Zhiling Luo, Shijian Li, Laurence T. Yang, Anind K. Dey, Gang Pan 0001 |
IEEE Trans. Ind. Informatics | 6 |
| 2020 | Forecasting Price Trend of Bulk Commodities Leveraging Cross-domain Open Data FusionabstractForecasting price trend of bulk commodities is important in international trade, not only for markets participants to schedule production and marketing plans but also for government administrators to adjust policies. Previous studies cannot support accurate fine-grained short-term prediction, since they mainly focus on coarse-grained long-term prediction using historical data. Recently, cross-domain open data provides possibilities to conduct fine-grained price forecasting, since they can be leveraged to extract various direct and indirect factors of the price. In this article, we predict the price trend over upcoming days, by leveraging cross-domain open data fusion. More specifically, we formulate the price trend into three classes (rise, slight-change, and fall), and then we predict the specific class in which the price trend of the future day lies. We take three factors into consideration: (1) supply factor considering sources providing bulk commodities,<?brk?> (2) demand factor focusing on vessel transportation with reflection of short time needs, and (3) expectation factor encompassing indirect features (e.g., air quality) with latent influences. A hybrid classification framework is proposed for the price trend forecasting. Evaluation conducted on nine real-world cross-domain open datasets shows that our framework can forecast the price trend accurately, outperforming multiple state-of-the-art baselines. Binbin Zhou 0005, Sha Zhao, Longbiao Chen, Shijian Li, Zhaohui Wu 0001, Gang Pan 0001 |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2019 | AppUsage2Vec: Modeling Smartphone App Usage for PredictionabstractApp usage prediction, i.e. which apps will be used next, is very useful for smartphone system optimization, such as operating system resource management, battery energy consumption optimization, and user experience improvement as well. However, it is still challenging to achieve usage prediction of high accuracy. In this paper, we propose a novel framework for app usage prediction, called AppUsage2Vec, inspired by Doc2Vec. It models app usage records by considering the contribution of different apps, user personalized characteristics, and temporal context. We measure the contribution of each app to the target app by introducing an app-attention mechanism. The user personalized characteristics in app usage are learned by a module of dual-DNN. Furthermore, we encode the top-k supervised information in loss function for training the model to predict the app most likely to be used next. The AppUsage2Vec was evaluated on a dataset of 10,360 users and 46,434,380 records in three months. The results demonstrate the state-of-the-art performance. Sha Zhao, Zhiling Luo, Ziwen Jiang, Shijian Li, Jianwei Yin, Gang Pan 0001 |
ICDE | 6 |
| 2019 | Investigating smartphone user differences in their application usage behaviors: an empirical study
Sha Zhao, Yizhi Xu, Xiaojuan Ma, Zhiling Luo, Shijian Li, Anind K. Dey, Gang Pan 0001 |
CCF Trans. Pervasive Comput. Interact. | 6 |
| 2019 | User profiling from their use of smartphone applications: A survey
Sha Zhao, Shijian Li, Julian Ramos 0001, Zhiling Luo, Ziwen Jiang, Anind K. Dey, Gang Pan 0001 |
Pervasive Mob. Comput. | 2 |
| 2016 | Scalable user assignment in power grids: a data driven approachabstractThe fast pace of global urbanization is drastically changing the population distributions over the world, which leads to significant changes in geographical population densities. Such changes in turn alter the underlying geographical power demand over time, and drive power substations to become over-supplied (demand << capacity) or under-supplied (demand ≈ capacity). In this paper, we make the first attempt to investigate the problem of power substation-user assignment by analyzing large-scale power grid data. We develop a Scalable Power User Assignment (SPUA) framework, that takes large-scale spatial power user/substation distribution data and temporal user power consumption data as input, and assigns users to substations, in a manner that minimizes the maximum substation utilization among all substations. To evaluate the performance of our SPUA framework, we conduct evaluations on real power consumption data and user/substation location data collected from a province in China for 35 days in 2015. The evaluation results demonstrate that our SPUA framework can achieve a 20%--65% reduction on the maximum substation utilization, and 2 to 3.7 times reduction on total transmission loss over other baseline methods. Bo Lyu, Shijian Li, Jie Fu 0002, Andrew C. Trapp, Haiyong Xie 0001, Yong Liao 0003 |
SIGSPATIAL/GIS | 2 |
| 2016 | Dynamic cluster-based over-demand prediction in bike sharing systemsabstractBike sharing is booming globally as a green transportation mode, but the occurrence of over-demand stations that have no bikes or docks available greatly affects user experiences. Directly predicting individual over-demand stations to carry out preventive measures is difficult, since the bike usage pattern of a station is highly dynamic and context dependent. In addition, the fact that bike usage pattern is affected not only by common contextual factors (e.g., time and weather) but also by opportunistic contextual factors (e.g., social and traffic events) poses a great challenge. To address these issues, we propose a dynamic cluster-based framework for over-demand prediction. Depending on the context, we construct a weighted correlation network to model the relationship among bike stations, and dynamically group neighboring stations with similar bike usage patterns into clusters. We then adopt Monte Carlo simulation to predict the over-demand probability of each cluster. Evaluation results using real-world data from New York City and Washington, D.C. show that our framework accurately predicts over-demand clusters and outperforms the baseline methods significantly. Longbiao Chen, Daqing Zhang 0001, Leye Wang, Dingqi Yang, Xiaojuan Ma, Shijian Li, Zhaohui Wu 0001, Gang Pan 0001, Thi Mai Trang Nguyen, Jérémie Jakubowicz |
UbiComp | 6 |
| 2016 | Discovering different kinds of smartphone users through their application usage behaviorsabstractUnderstanding smartphone users is fundamental for creating better smartphones, and improving the smartphone usage experience and generating generalizable and reproducible research. However, smartphone manufacturers and most of the mobile computing research community make a simplifying assumption that all smartphone users are similar or, at best, constitute a small number of user types, based on their behaviors. Manufacturers design phones for the broadest audience and hope they work for all users. Researchers mostly analyze data from smartphone-based user studies and report results without accounting for the many different groups of people that make up the user base of smartphones. In this work, we challenge these elementary characterizations of smartphone users and show evidence of the existence of a much more diverse set of users. We analyzed one month of application usage from 106,762 Android users and discovered 382 distinct types of users based on their application usage behaviors, using our own two-step clustering and feature ranking selection approach. Our results have profound implications on the reproducibility and reliability of mobile computing studies, design and development of applications, determination of which apps should be pre-installed on a smartphone and, in general, on the smartphone usage experience for different types of users. Sha Zhao, Julian Ramos 0001, Jianrong Tao, Ziwen Jiang, Shijian Li, Zhaohui Wu 0001, Gang Pan 0001, Anind K. Dey |
UbiComp | 5 |
| 2016 | Container Port Performance Measurement and Comparison Leveraging Ship GPS Traces and Maritime Open DataabstractContainer ports are generally measured and compared using performance indicators such as container throughput and facility productivity. Being able to measure the performance of container ports quantitatively is of great importance for researchers to design models for port operation and container logistics. Instead of relying on the manually collected statistical information from different port authorities and shipping companies, we propose to leverage the pervasive ship GPS traces and maritime open data to derive port performance indicators, including ship traffic, container throughput, berth utilization, and terminal productivity. These performance indicators are found to be directly related to the number of container ships arriving at the terminals and the number of containers handled at each ship. Therefore, we propose a framework that takes the ships' container-handling events at terminals as the basis for port performance measurement. With the inferred port performance indicators, we further compare the strengths and weaknesses of different container ports at the terminal level, port level, and region level, which can potentially benefit terminal productivity improvement, liner schedule optimization, and regional economic development planning. In order to evaluate the proposed framework, we conduct extensive studies on large-scale real-world GPS traces of container ships collected from major container ports worldwide through the year, as well as various maritime open data sources concerning ships and ports. Evaluation results confirm that the proposed framework not only can accurately estimate various port performance indicators but also effectively produces port comparison results such as port performance ranking and port region comparison. Longbiao Chen, Daqing Zhang 0001, Xiaojuan Ma, Leye Wang, Shijian Li, Zhaohui Wu 0001, Gang Pan 0001 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2015 | Bike sharing station placement leveraging heterogeneous urban open dataabstractBike sharing systems have been deployed in many cities to promote green transportation and a healthy lifestyle. One of the key factors for maximizing the utility of such systems is placing bike stations at locations that can best meet users' trip demand. Traditionally, urban planners rely on dedicated surveys to understand the local bike trip demand, which is costly in time and labor, especially when they need to compare many possible places. In this paper, we formulate the bike station placement issue as a bike trip demand prediction problem. We propose a semi-supervised feature selection method to extract customized features from the highly variant, heterogeneous urban open data to predict bike trip demand. Evaluation using real-world open data from Washington, D.C. and Hangzhou shows that our method can be applied to different cities to effectively recommend places with higher potential bike trip demand for placing future bike stations. Longbiao Chen, Daqing Zhang 0001, Gang Pan 0001, Xiaojuan Ma, Dingqi Yang, Kostadin Kushlev, Wangsheng Zhang, Shijian Li |
UbiComp | 8 |
| 2015 | City-Scale Social Event Detection and Evaluation with Taxi TracesabstractA social event is an occurrence that involves lots of people and is accompanied by an obvious rise in human flow. Analysis of social events has real-world importance because events bring about impacts on many aspects of city life. Traditionally, detection and impact measurement of social events rely on social investigation, which involves considerable human effort. Recently, by analyzing messages in social networks, researchers can also detect and evaluate country-scale events. Nevertheless, the analysis of city-scale events has not been explored. In this article, we use human flow dynamics, which reflect the social activeness of a region, to detect social events and measure their impacts. We first extract human flow dynamics from taxi traces. Second, we propose a method that can not only discover the happening time and venue of events from abnormal social activeness, but also measure the scale of events through changes in such activeness. Third, we extract traffic congestion information from traces and use its change during social events to measure their impact. The results of experiments validate the effectiveness of both the event detection and impact measurement methods. Wangsheng Zhang, Guande Qi, Gang Pan 0001, Hua Lu 0001, Shijian Li, Zhaohui Wu 0001 |
ACM Trans. Intell. Syst. Technol. | 5 |
| 2015 | Understanding Taxi Service Strategies From Taxi GPS TracesabstractTaxi service strategies, as the crowd intelligence of massive taxi drivers, are hidden in their historical time-stamped GPS traces. Mining GPS traces to understand the service strategies of skilled taxi drivers can benefit the drivers themselves, passengers, and city planners in a number of ways. This paper intends to uncover the efficient and inefficient taxi service strategies based on a large-scale GPS historical database of approximately 7600 taxis over one year in a city in China. First, we separate the GPS traces of individual taxi drivers and link them with the revenue generated. Second, we investigate the taxi service strategies from three perspectives, namely, passenger-searching strategies, passenger-delivery strategies, and service-region preference. Finally, we represent the taxi service strategies with a feature matrix and evaluate the correlation between service strategies and revenue, informing which strategies are efficient or inefficient. We predict the revenue of taxi drivers based on their strategies and achieve a prediction residual as less as 2.35 RMB/h,1which demonstrates that the extracted taxi service strategies with our proposed approach well characterize the driving behavior and performance of taxi drivers. Daqing Zhang 0001, Lin Sun 0009, Bin Li 0015, Chao Chen 0004, Gang Pan 0001, Shijian Li, Zhaohui Wu 0001 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2014 | Container throughput estimation leveraging ship GPS traces and open dataabstractTraditionally, the port container throughput, a crucial measurement of regional economic development, was manually collected by port authorities. This requires a large amount of human effort and often delays publication of this important figure. In this paper, by leveraging ubiquitous positioning techniques and open data, we propose a two-phase approach to estimation of port container throughput in real-time. First, we obtain the number of container ships arriving at berth by analyzing the ships' GPS traces. Then we estimate the throughput of each ship, in terms of number of containers transshipped, by considering the ship's berthing time, capacity, length, breadth, and crane operation performance, as extracted from different data sources. Evaluation results using real-world datasets from Hong Kong and Singapore show that the proposed approach not only estimates the container throughput quite accurately, but also outperforms the baseline method significantly. Longbiao Chen, Daqing Zhang 0001, Gang Pan 0001, Leye Wang, Xiaojuan Ma, Chao Chen 0004, Shijian Li |
UbiComp | 7 |
| 2013 | Online Community Detection for Large Complex Networks
Wangsheng Zhang, Gang Pan 0001, Zhaohui Wu 0001, Shijian Li |
IJCAI | 4 |
| 2013 | B-Planner: Night bus route planning using large-scale taxi GPS tracesabstractTaxi GPS traces provide us with rich information about the human mobility pattern in modern cities. Instead of designing the bus route based on inaccurate human survey regarding people's mobility pattern, we intend to address the night-bus route planning issue by leveraging taxi GPS traces. In this paper, we propose a two-phase approach based on the crowd-sourced GPS data for night-bus route planning. In the first phase, we develop a process to cluster “hot” areas with dense passenger pick-up/drop-off, and then propose effective methods to split big “hot” areas into clusters and identify a location in each cluster as a candidate bus stop. In the second phase, given the bus route origin, destination, candidate bus stops as well as bus operation time constraints, we derive several effective rules to build bus routing graph and prune the invalid stops and edges iteratively. We further develop two heuristic algorithms to automatically generate candidate bus routes, and finally we select the best route which expects the maximum number of passengers under the given conditions. To validate the effectiveness of the proposed approach, extensive empirical studies are performed on a real-world taxi GPS data set which contains more than 1.57 million passenger delivery trips, generated by 7,600 taxis for a month in Hangzhou, China. Chao Chen 0004, Daqing Zhang 0001, Zhi-Hua Zhou, Nan Li 0019, Tülin Atmaca, Shijian Li |
PerCom | 6 |
| 2013 | Real Time Anomalous Trajectory Detection and Analysis
Lin Sun 0009, Daqing Zhang 0001, Chao Chen 0004, Pablo Samuel Castro, Shijian Li, Zonghui Wang |
Mob. Networks Appl. | 5 |
| 2013 | iBOAT: Isolation-Based Online Anomalous Trajectory DetectionabstractTrajectories obtained from Global Position System (GPS)-enabled taxis grant us an opportunity not only to extract meaningful statistics, dynamics, and behaviors about certain urban road users but also to monitor adverse and/or malicious events. In this paper, we focus on the problem of detecting anomalous routes by comparing the latter against time-dependent historically “normal” routes. We propose an online method that is able to detect anomalous trajectories “on-the-fly” and to identify which parts of the trajectory are responsible for its anomalousness. Furthermore, we perform an in-depth analysis on around 43 800 anomalous trajectories that are detected out from the trajectories of 7600 taxis for a month, revealing that most of the anomalous trips are the result of conscious decisions of greedy taxi drivers to commit fraud. We evaluate our proposed isolation-based online anomalous trajectory (iBOAT) through extensive experiments on large-scale taxi data, and it shows that iBOAT achieves state-of-the-art performance, with a remarkable performance of the area under a curve (AUC)$\geq$0.99. Chao Chen 0004, Daqing Zhang 0001, Pablo Samuel Castro, Nan Li 0019, Lin Sun 0009, Shijian Li, Zonghui Wang |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2013 | Land-Use Classification Using Taxi GPS TracesabstractDetailed land use, which is difficult to obtain, is an integral part of urban planning. Currently, GPS traces of vehicles are becoming readily available. It conveys human mobility and activity information, which can be closely related to the land use of a region. This paper discusses the potential use of taxi traces for urban land-use classification, particularly for recognizing the social function of urban land by using one year's trace data from 4000 taxis. First, we found that pick-up/set-down dynamics, extracted from taxi traces, exhibited clear patterns corresponding to the land-use classes of these regions. Second, with six features designed to characterize the pick-up/set-down pattern, land-use classes of regions could be recognized. Classification results using the best combination of features achieved a recognition accuracy of 95%. Third, the classification results also highlighted regions that changed land-use class from one to another, and such land-use class transition dynamics of regions revealed unusual real-world social events. Moreover, the pick-up/set-down dynamics could further reflect to what extent each region is used as a certain class. Gang Pan 0001, Guande Qi, Zhaohui Wu 0001, Daqing Zhang 0001, Shijian Li |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2012 | FlyingBuddy2: a brain-controlled assistant for the handicappedabstractThe motor impaired people have much limit in moving. The devices augmenting their mobility will be much helpful for improving their living experiences. This poster develops a brain-controlled assistive system, called FlyingBuddy2, to aid the handicapped in mobility. It uses the brain EEG signals to directly control a quadrotor. Signals from an EEG headset are transmitted wirelessly to a computer, then the decoded brain signals are converted to trigger the quadrotor to move in 3D space. Three applications are developed: thinking to play games, thinking to see, and thinking to take pictures. Yipeng Yu, Weidong Hua, Shijian Li, Yueming Wang 0001, Gang Pan 0001 |
UbiComp | 4 |
| 2012 | Mining the semantics of origin-destination flows using taxi tracesabstractOrigin-destination(OD) flows reflect both human activity and urban dynamic in a city. However, our understanding about their patterns remains limited. In this paper, we study the GPS traces of taxis in a city with several millions people, China and find that there are significant patterns under the OD flows constructed from taxis' random motion. Our spatiotemporal analysis shows that those patterns have close relationship with the semantics of OD flows, hence we can mine the semantics of OD flows from raw GPS trace data. The approach we proposed offers a novel way to explore the human mobility and location characteristic. Wangsheng Zhang, Shijian Li, Gang Pan 0001 |
UbiComp | 2 |
| 2012 | SmartShadow-K: an practical knowledge network for joint context inference in everyday lifeabstractSmart environments require to percept conditions of people. Current context-aware systems mainly model limited user situations, which constrains their coverage and effect in real world usage. This paper proposes an encyclopedic knowledge network to enable practical context inference in our daily life by: 1) expressing essential semantics of contextual concepts and relations into a well-informed relational network, and 2) exploiting relational semantics to infer various contexts simultaneously. The performance of the approach is validated in real challenging problems and compared with inference of human being. Li Zhang 0045, Gang Pan 0001, Zhaohui Wu 0001, Shijian Li, Cho-Li Wang |
UbiComp | 4 |
| 2012 | Prediction of urban human mobility using large-scale taxi traces and its applications
Gang Pan 0001, Zhaohui Wu 0001, Guande Qi, Shijian Li, Daqing Zhang 0001, Wangsheng Zhang, Zonghui Wang |
Frontiers Comput. Sci. China | 5 |
| 2011 | Tilt & touch: mobile phone for 3D interactionabstractMobile phones are becoming de facto pervasive devices for people's daily use. This demonstration illustrates a new interaction, Tilt & Touch, to enable a smart phone to be a 3D controller. It exploits capacitive touchscreen and built-in MEMS motion sensors. When people want to navigate in a virtual reality environment on a large display, they can tilt the phone for viewpoint transforming, touch the phone screen for avatar moving, and pinch screen for viewing camera zooming. The virtual objects in the virtual reality environment can be rotated accordingly by tilting the phone. Yuan Du, Haoyi Ren, Gang Pan 0001, Shijian Li |
UbiComp | 4 |
| 2011 | FlyingBuddy: augment human mobility and perceptibilityabstractTechnologies keep evolving to strengthen and further people's abilities in many aspects. For instance, vehicles expand the range of human moving while mobile phones boost the range of human communication. In this video, we develop a novel mini unmanned aerial vehicle (mini-UAV) named FlyingBuddy to augment human mobility and perceptibility. This prototype is made up of off the shelf components AR. Drone and iPhones with customized software. With help of the built-in magnetometer, GPS, and cameras, as well as Bluetooth, Wi-Fi and 3G connectivity, FlyingBuddy is capable of both manual controlled and self-piloted flying. It provides four typical services: flying to buy, flying to see, flying to report accident, and flying to take pictures. Haoyi Ren, Weidong Hua, Gang Pan 0001, Shijian Li, Zhaohui Wu 0001 |
UbiComp | 5 |
| 2011 | iBAT: detecting anomalous taxi trajectories from GPS tracesabstractGPS-equipped taxis can be viewed as pervasive sensors and the large-scale digital traces produced allow us to reveal many hidden "facts" about the city dynamics and human behaviors. In this paper, we aim to discover anomalous driving patterns from taxi's GPS traces, targeting applications like automatically detecting taxi driving frauds or road network change in modern cites. To achieve the objective, firstly we group all the taxi trajectories crossing the same source destination cell-pair and represent each taxi trajectory as a sequence of symbols. Secondly, we propose an Isolation-Based Anomalous Trajectory (iBAT) detection method and verify with large scale taxi data that iBAT achieves remarkable performance (AUC>0.99, over 90% detection rate at false alarm rate of less than 2%). Finally, we demonstrate the potential of iBAT in enabling innovative applications by using it for taxi driving fraud detection and road network change detection. Daqing Zhang 0001, Nan Li 0019, Zhi-Hua Zhou, Chao Chen 0004, Lin Sun 0009, Shijian Li |
UbiComp | 6 |
| 2011 | Real-Time Detection of Anomalous Taxi Trajectories from GPS Traces
Chao Chen 0004, Daqing Zhang 0001, Pablo Samuel Castro, Nan Li 0019, Lin Sun 0009, Shijian Li |
MobiQuitous | 6 |
| 2010 | Semantic Device Bus for Internet of ThingsabstractThe vision of the Internet of things is very appealing, and gains more and more attention. Since mobile devices in the Internet become more complex and heterogeneous, device collaboration will be full of technical challenges. How to integrate different systems and heterogeneous devices is a big problem. In order to overcome the problem, it is very important to describe and match the heterogeneous device services with semantics. We use web service interface specification to wrap device, and OWL to describe services in semantic. In this paper we introduce a semantic device bus for Internet of things, which will allow service creation of different devices, management of device services, semantic description and matching of device services, and complex collaboration of device services. The bus provides a fundamental platform for large-scale applications of the Internet of things. Shijian Li, Li Zhang 0045, Gang Pan 0001 |
EUC | 2 |
| 2010 | Modeling Files with Context Streams
Qunjie Qiu, Gang Pan 0001, Shijian Li |
UIC | 3 |
| 2010 | Activity Recognition on an Accelerometer Embedded Mobile Phone with Varying Positions and Orientations
Lin Sun 0009, Daqing Zhang 0001, Bin Li 0015, Bin Guo 0001, Shijian Li |
UIC | 5 |
| 2010 | GeeAir: a universal multimodal remote control device for home appliances
Gang Pan 0001, Daqing Zhang 0001, Zhaohui Wu 0001, Yingchun Yang, Shijian Li |
Pers. Ubiquitous Comput. | 6 |
| 2009 | SmartShadow: Modeling A User-centric Mobile Virtual SpaceabstractThis paper attempts to model pervasive computing environments as a user-centric ldquoSmartShadowrdquo using the BDP (belief-desire-plan) user model, which maps pervasive computing environments into a dynamic virtual user space. SmartShadow will follow the user to provide him with pervasive services, just like his shadow in the physical world. In the BDP model, desires of a user are inferred from his belief set, and plans are made to satisfy each desire. Pervasive service is introduced to describe computing resources in the cyberspace, which can be organized by the user's BDP to accomplish his desires. The composition process maps pervasive services into a user's SmartShadow. The model is logically natural and simple, and can flexibly model dynamics of pervasive computing spaces. In addition, we implement a simulation system to verify and evaluate the SmartShadow model. Li Zhang 0045, Gang Pan 0001, Zhaohui Wu 0001, Shijian Li, Cho-Li Wang |
PerCom | 4 |
| 2009 | Gesture Recognition with a 3-D Accelerometer
Gang Pan 0001, Daqing Zhang 0001, Guande Qi, Shijian Li |
UIC | 5 |