Zhongmin Cai

dblp:05/7884 · DBLP profile ↗
← Back
52ranked-venue papers
2as first author
19since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 1 first-author · 8 since 2021Security and privacy · 14 · 1 since 2021Human-computer interaction and ubiquitous computing · 11 · 1 first-author · 6 since 2021Computer networks · 5 · 1 since 2021Databases, data management, data science and information retrieval · 5 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 since 2021Systems, architecture and hardware · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 DroidRetriever: A Transparent and Steerable Automation System for Collaborative Mobile Information Seeking
abstract
Information seeking on mobile devices is often fragmented, trapping users in repetitive cycles of context switching and data re-entry, which increases cognitive load and disrupts workflow. Existing mobile agents provide limited cross-source integration and are largely opaque, presenting progress as a linear feed with few opportunities to intervene, steer, or take control. We present DroidRetriever, a transparent, steerable system for cross-source mobile information seeking. It accepts voice or typed input and the multi-LLM system decomposes the task, navigates to target pages, takes screenshots, and synthesizes a concise report with citation-linked screenshots. We make the process transparent through a progress dashboard combining sub-task progress and real-time exploration maps for seamless takeover. DroidRetriever also pauses on detected privacy or high-risk screens and prompts intervention. Across 35 tasks over 24 apps, experiments and user studies demonstrate improvements in coverage, transparency, and reduced workload. We release our code at https://github.com/AkimotoAyako/DroidRetriever.
Yiheng Bian, Yunpeng Song, Guiyu Ma, Rongrong Zhu, Zhongmin Cai
CHI5
2026 AReframedChair: Reframing the Empty Chair through Dyadic and Triadic AR-Mediated Self-Embodiment
abstract
Immersive technologies are increasingly applied in therapeutic and well-being practices, yet most AR systems focus on dyadic client–avatar interactions and overlook richer therapeutic structures that involve therapists. We introduce AReframedChair, an AR system that reimagines the traditional Empty Chair technique by enabling self-dialogue with a personalized avatar representing one’s past or future self. In a between-subjects study with 60 adults, we compared the traditional Empty Chair method with two AR-reframed modes: Dyadic (client–avatar) and Triadic (client–avatar– therapist). Participants’ survey responses showed that the Dyadic mode elicited greater positive affect and self-compassion in the past-self scenarios, whereas the Triadic mode produced stronger gains in motivation and reflections in future-self scenarios. Thematic analysis further revealed distinct roles: the Avatar facilitated emotional entry, reassurance, and cognitive reframing, while the Therapists intervened at critical moments to down-regulate intensity, redirect attention, and enhance reflection. These findings open up new design pathways for mental health technologies.
Ling Ling, Yunpeng Song, Yun Huang 0003, Zhongmin Cai
CHI5
2026 A Multi-Stage Structural Captioning Framework for Enhancing Chinese Image-to-Video Generation in Baidu
abstract
The rapid advancement of video generation technology, particularly in the domain of image-to-video generation, is significantly transforming both personal and industrial applications. The performance of video generation models is heavily influenced by the quality of data, specifically video content and captions. Among these, the quality of video captions directly impacts the model's ability to follow user instructions. Currently, video captions are typically generated using vision-language model (VLM)-based methods. However, issues such as hallucinations, incomplete descriptions, and inaccuracies remain prevalent. To address these challenges, we propose a novel optimization framework for a multi-expert, structured video captioning model, specifically designed to enhance the performance of Chinese image-to-video tasks in Baidu's business scenarios. First, we construct a structured video caption dataset by combining a large language model (LLM) with human annotations, and introduce a hallucination-aware adaptive GRPO algorithm. A two-stage fine-tuning alignment strategy is employed to improve the accuracy of fine-grained content descriptions and mitigate hallucinations. Second, we design independent video sub-expert models to characterize subject motion intensity and camera movement, thereby expanding the dimensionality of video captions and addressing the issue of inaccurate sub-dimension descriptions in a single VLM. Finally, we leverage multi-level structured video captions to guide the training of the video generation model, resulting in improved quality and consistency in video generation. Extensive experiments demonstrate that our method consistently outperforms the baseline. Specifically, it achieves a 2.49% improvement in F1-score on event-level cross-validation and reduces hallucination occurrences by 30.9%. Moreover, our framework significantly enhances the quality of Chinese video generation, yielding a 2.9% increase in the VBench total score and a 9.2% improvement in the proportion of high-quality generated videos.
Zhipeng Jin, Xiawei Li, Wen Tao, Yi Yang 0031, Cong Han 0002, Shuanglong Li, Zhongmin Cai
KDD (1)9
2026 Reasoning step by step via a neural-symbolic geometry problem solver
Yaxian Wang, Bifan Wei, Yinghong Ma, Xudong Jiang 0001, Henghui Ding, Zhongmin Cai, Jun Liu 0002
Pattern Recognit.6
2025 Predicting User Behavior in Smart Spaces with LLM-Enhanced Logs and Personalized Prompts
abstract
Enhancing the intelligence of smart systems, such as smart homes, smart vehicles, and smart grids, critically depends on developing sophisticated planning capabilities that can anticipate the next desired function based on historical interactions. While existing methods view user behaviors as sequential data and apply models like RNNs and Transformers to predict future actions, they often fail to incorporate domain knowledge and capture personalized user preferences. In this paper, we propose a novel approach that incorporates LLM-enhanced logs and personalized prompts. Our approach first constructs a graph that captures individual behavior preferences derived from their interaction histories. This graph effectively transforms into a soft continuous prompt that precedes the sequence of user behaviors. Then our approach leverages the vast general knowledge and robust reasoning capabilities of a pretrained LLM to enrich the oversimplified and incomplete log records. By enhancing these logs semantically, our approach better understands the user's actions and intentions, especially for those rare events in the dataset. We evaluate the method across four real-world datasets from both smart vehicle and smart home settings. The findings validate the effectiveness of our LLM-enhanced description and personalized prompt, shedding light on potential ways to advance the intelligence of smart space.
Yunpeng Song, Yiheng Bian, Zhongmin Cai
AAAI4
2025 EyeSee: Enhancing Art Appreciation through Anthropomorphic Interpretations from Multiple Perspectives
Hangyue Zhang, Andrea Yaoyun Cui, Zisong Ma, Yunpeng Song, Zhongmin Cai, Yun Huang 0003
CHI6
2025 Retrieval-Augmented Image Captioning and Generation with Entity Concepts Enhancement for Baidu Multimodal Advertising
abstract
Recent advancements in generative artificial intelligence are driving a significant transformation in information retrieval and content generation, creating substantial opportunities for online advertising. Text-to-image generation technology has become increasingly prevalent in advertising content production, demonstrating promising performance improvements in terms of semantic relevance and visual appeal. However, existing models often suffer from inadequate representation of entity concepts, such as prominent product brands and recognizable landmarks. This inherent limitation subsequently leads to notable deficiencies in brand tonality, industry-specific relevance, and market adaptability of the generated advertising content. To address this challenge, we propose a multimodal ad content generation framework specifically engineered for online advertising system, particularly focused on resolving the deficiency in entity concepts. Our framework is comprised of two phases: first, an image captioning module with entity-aware learning based on multimodal large language model, leveraging retrieval-augmented techniques to incorporate entity concepts into image descriptions; second, a text-to-image diffusion model refined on image-text pairs enriched with entity concepts to facilitate entity-grounded image generation. Extensive experiments validate the effectiveness of our framework, demonstrating superior performance in both image captioning and image generation compared to existing methods, particularly in the accuracy of depiction of relevant entities in advertising images. Moreover, the deployment of the framework in the system primary traffic of Baidu Search Ads, has brought significant enhancements to advertisement revenue for both advertisers and the platform.
Kang Zhao 0002, Zhipeng Jin, Wen Tao, Yi Yang 0031, Cong Han 0002, Shuanglong Li, Zhongmin Cai
SIGIR8
2025 Quantitative Estimation of Human Height and Weight Using Motion Data From Multiple Smart Devices
abstract
This article proposes a methodological framework for quantitative estimations of height and weight using behavioral data collected from smart devices. We analyze the connections between height and weight information and behavioral data from three aspects: walking speed, stride length, and step frequency and then extract two kinds of motion features including basic kinematic features and advanced features which use statistical measurements summarizing the dynamics of walking behavior over time and relative intensity of walking speed change, energy cost of one step during walking, and walking frequency, respectively, to describe the motion behavior. After that, we qualitatively and quantitatively analyze the complementarity of different motion data sources and show that more useful information existed in multisource motion data than that of only one motion data source. Based on this, we propose a feature fusion approach named Serial+CC to dealing with the relationships between all motion features from multiple smart devices and user traits of height and weight and then a fused feature set with high discrimination and low complexity is constructed. Finally, five regression models of SVM, BP neural networks, Random Forest, LSTM, and BiLSTM are built with the fused feature set. Empirical evaluations were performed on a dataset collected from 56 subjects. The results demonstrate that motion data collected from smart devices can be used for height and weight quantitative estimation. The results also illustrate our method of using motion data collected from multiple smart devices can achieve better performance than those of only using one smart devices. The best performance is achieved with average errors of 0.95% (1.59 cm) and 4.75% (2.90 kg) for height and weight estimations, respectively, in the scenario of using multiple devices.
Jianmin Dong 0001, Zhongmin Cai
IEEE Trans. Comput. Soc. Syst.2
2025 Bilevel Optimized Collusion Attacks Against Gait Recognizer
abstract
Extensive investigations have revealed that the gait recognition system is always vulnerable to impersonation attacks, which pose significant threats to the identity access security. Previous impersonation strategies have primarily focused on mimicking the victim’s walking style or probing the similar gait features to merely manipulate the input samples, without concurrently undermining the built-in model of the gait recognizer, thereby failing to achieve cost-effective attacks. In contrast to these existing heuristic approaches, we propose an optimal adversarial complicity strategy, called collusion attack, which leverages the tight collaboration between an external attacker and an internal spy to tie up into the close colluder, simultaneously enabling the input-&model-corrupted tampering modes and misleading the gait recognizer more powerfully and stealthily for misidentifying the illegitimate Alice as legitimate Bob. Specifically, we formulate a bilevel optimization problem to model such a leader-follower Stackelberg game with sequentially adversarial interaction process between the colluders and gait recognizer. Further, to solve this challenging bilevel problem efficiently, we absorb the Lagrangian dual theory and linearization representation method to reformulate a tractable mixed integer program. Finally, we perform comparison and ablation experiments with the state-of-the-art attack modes on single-&multi-source gait datasets to verify the validity of our collusion strategy in inducing the mistaken identity with great success rate, high confidence, and low cost. Empirical results also shed light on key insights in mitigating the collusion attacks and enhancing the gait recognition robustness to safeguard the identity access applications.
Jianmin Dong 0001, Datian Peng, Zhongmin Cai, Bo Zeng 0001
IEEE Trans. Inf. Forensics Secur.3
2024 Towards Building Condition-Based Cross-Modality Intention-Aware Human-AI Cooperation under VR Environment
abstract
To address critical challenges in effectively identifying user intent and forming relevant information presentations and recommendations in VR environments, we propose an innovative condition-based multi-modal human-AI cooperation framework. It highlights the intent tuples (intent, condition, intent prompt, action prompt) and 2-Large-Language-Models (2-LLMs) architecture. This design, utilizes “condition” as the core to describe tasks, dynamically match user interactions with intentions, and empower generations of various tailored multi-modal AI responses. The architecture of 2-LLMs separates the roles of intent detection and action generation, decreasing the prompt length and helping with generating appropriate responses. We implemented a VR-based intelligent furniture purchasing system based on the proposed framework and conducted a three-phase comparative user study. The results conclusively demonstrate the system’s superiority in time efficiency and accuracy, intention conveyance improvements, effective product acquisitions, and user satisfaction and cooperation preference. Our framework provides a promising approach towards personalized and efficient user experiences in VR.
Ziyao He, Yunpeng Song, Zhongmin Cai
CHI4
2024 QGEval: Benchmarking Multi-dimensional Evaluation for Question Generation
abstract
Automatically generated questions often suffer from problems such as unclear expression or factual inaccuracies, requiring a reliable and comprehensive evaluation of their quality.Human evaluation is widely used in the field of question generation (QG) and serves as the gold standard for automatic metrics.However, there is a lack of unified human evaluation criteria, which hampers consistent and reliable evaluations of both QG models and automatic metrics.To address this, we propose QGEval, a multi-dimensional Evaluation benchmark for Question Generation, which evaluates both generated questions and existing automatic metrics across 7 dimensions: fluency, clarity, conciseness, relevance, consistency, answerability, and answer consistency.We demonstrate the appropriateness of these dimensions by examining their correlations and distinctions.Through consistent evaluations of QG models and automatic metrics with QGEval, we find that 1) most QG models perform unsatisfactorily in terms of answerability and answer consistency, and 2) existing metrics fail to align well with human judgments when evaluating generated questions across the 7 dimensions.We expect this work to foster the development of both QG technologies and their evaluation.
Weiping Fu, Bifan Wei, Jianxiang Hu, Zhongmin Cai, Jun Liu 0002
EMNLP4
2024 Beyond Single Stationary Policies: Meta-Task Players as Naturally Superior Collaborators
abstract
In human-AI collaborative tasks, the distribution of human behavior, influenced by mental models, is non-stationary, manifesting in various levels of initiative and different collaborative strategies. A significant challenge in human-AI collaboration is determining how to collaborate effectively with humans exhibiting non-stationary dynamics. Current collaborative agents involve initially running self-play (SP) multiple times to build a policy pool, followed by training the final adaptive policy against this pool. These agents themselves are a single policy network, which is $\textbf{insufficient for handling non-stationary human dynamics}$. We discern that despite the inherent diversity in human behaviors, the $\textbf{underlying meta-tasks within specific collaborative contexts tend to be strikingly similar}$. Accordingly, we propose $\textbf{C}$ollaborative $\textbf{B}$ayesian $\textbf{P}$olicy $\textbf{R}$euse ($\textbf{CBPR}$), a novel Bayesian-based framework that $\textbf{adaptively selects optimal collaborative policies matching the current meta-task from multiple policy networks}$ instead of just selecting actions relying on a single policy network. We provide theoretical guarantees for CBPR's rapid convergence to the optimal policy once human partners alter their policies. This framework shifts from directly modeling human behavior to identifying various meta-tasks that support human decision-making and training meta-task playing (MTP) agents tailored to enhance collaboration. Our method undergoes rigorous testing in a well-recognized collaborative cooking simulator, $\textit{Overcooked}$. Both empirical results and user studies demonstrate CBPR's superior competitiveness compared to existing baselines.
Zhaoming Tian, Yunpeng Song, Xiangliang Zhang 0001, Zhongmin Cai
NeurIPS5
2024 VisionTasker: Mobile Task Automation Using Vision Based UI Understanding and LLM Task Planning
abstract
Mobile task automation is an emerging field that leverages AI to streamline and optimize the execution of routine tasks on mobile devices, thereby enhancing efficiency and productivity. Traditional methods, such as Programming By Demonstration (PBD), are limited due to their dependence on predefined tasks and susceptibility to app updates. Recent advancements have utilized the view hierarchy to collect UI information and employed Large Language Models (LLM) to enhance task automation. However, view hierarchies have accessibility issues and face potential problems like missing object descriptions or misaligned structures. This paper introduces VisionTasker, a two-stage framework combining vision-based UI understanding and LLM task planning, for mobile task automation in a step-by-step manner. VisionTasker firstly converts a UI screenshot into natural language interpretations using a vision-based UI understanding approach, eliminating the need for view hierarchies. Secondly, it adopts a step-by-step task planning method, presenting one interface at a time to the LLM. The LLM then identifies relevant elements within the interface and determines the next action, enhancing accuracy and practicality. Extensive experiments show that VisionTasker outperforms previous methods, providing effective UI representations across four datasets. Additionally, in automating 147 real-world tasks on an Android smartphone, VisionTasker demonstrates advantages over humans in tasks where humans show unfamiliarity and shows significant improvements when integrated with the PBD mechanism. VisionTasker is open-source and available at https://github.com/AkimotoAyako/VisionTasker.
Yunpeng Song, Yiheng Bian, Yongtao Tang, Guiyu Ma, Zhongmin Cai
UIST5
2024 Performance evaluation of lightweight network-based bot detection using mouse movements
Hongfeng Niu, Yuxun Zhou, Jiading Chen, Zhongmin Cai
Eng. Appl. Artif. Intell.4
2024 Touch Authentication for Sharing Context Using Within-Group Similarity Structure
abstract
Sharing digital resources is a common practice in both work and personal life. Yet, sharing identical credentials, such as passwords or physical cards, not only poses significant security risks but also falls short in addressing the specific requirements of small local groups, such as parental controls, tracking user modifications, and easily updating access. To address this, we suggest a touch behavior-based method tailored for sharing in small local groups, designed to balance between ensuring relaxed security and maintaining practical functionality. Our approach aims to concurrently identify in-group users and detect out-of-group imposters. Specifically, our approach extracts effective identity representations that are robust to in-group variability and out-of-group uncertainty by learning a pair of touch-behavioral and within-group similarity embeddings. While the former captures the unique features of user touch characteristics, the latter reflects the typical group-wide similarity structure that an in-group user is expected to possess from a holistic perspective. Experimental results showcase the effectiveness of our method even with few samples for training. It maintains accuracy despite the group growing larger and shows resilience against the advanced attacks. This offers a promising way to keep group access both user-friendly and relatively secure, striking a crucial balance for small groups’ needs.
Yunpeng Song, Zhongmin Cai, Zhou Su 0001
IEEE Internet Things J.3
2023 Interaction of Thoughts: Towards Mediating Task Assignment in Human-AI Cooperation with a Capability-Aware Shared Mental Model
abstract
The existing work on task assignment of human-AI cooperation did not consider the differences between individual team members regarding their capabilities, leading to sub-optimal task completion results. In this work, we propose a capability-aware shared mental model (CASMM) with the components of task grouping and negotiation, which utilize tuples to break down tasks into sets of scenarios relating to difficulties and then dynamically merge the task grouping ideas raised by human and AI through negotiation. We implement a prototype system and a 3-phase user study for the proof of concept via an image labeling task. The result shows building CASMM boosts the accuracy and time efficiency significantly through forming the task assignment close to real capabilities within few iterations. It helps users better understand the capability of AI and themselves. Our method has the potential to generalize to other scenarios such as medical diagnoses and automatic driving in facilitating better human-AI cooperation.
Ziyao He, Yunpeng Song, Shurui Zhou, Zhongmin Cai
CHI4
2023 Exploring visual representations of computer mouse movements for bot detection using deep learning approaches
Hongfeng Niu, Ang Wei, Yunpeng Song, Zhongmin Cai
Expert Syst. Appl.4
2022 Reading Personality Preferences From Motion Patterns in Computer Mouse Operations
abstract
Personality not only plays essential roles in people's real lives, but also becomes an important factor for various online services. Traditional approaches to personality assessment usually ask users to answer a long list of questions and are thus not practical in many online systems. It is more desirable to use some universally available data in the system to perform personality assessment. This article reports a controlled study to investigate common mouse operations as a potential new type of data source for online personality assessment. We establish an elaborate personality-mouse behavior dataset from 146 subjects and propose kinematic and adjustment features to characterize mouse motion patterns. Statistical approaches and machine learning algorithms are employed to examine the connections between personality preferences and mouse motion features via correlation analysis, discernibility analysis, and personality recognition experiments. The results reveal some interesting expressions of cognitive personality perspectives reflected in mouse motion patterns, such asfast starting accelerationandbetter controller. Performance evaluation shows it is possible to recognize different personality preferences using mouse motion features with accuracies ranging from 60.6 to 78.3 percent. Our findings suggest a potential to use mouse operational behaviors as a new data source for personality assessment in various information systems.
Yinghui Zhao, Danmin Miao, Zhongmin Cai
IEEE Trans. Affect. Comput.3
2021 Learning Fundamental Visual Concepts Based on Evolved Multi-Edge Concept Graph
abstract
In general, visual media comprises a set of elements of basic semantics, named fundamental visual concepts, that may not be semantically decomposed, such as objects, scenes and actions. This paper proposes a dynamic learning framework for fundamental visual concept learning from image-textual description paired data based on an evolved multi-edge concept graph (EMCG). First, we construct a multi-edge concept graph to represent the relationships between visual concept instances, in which we introduce two types of edges named visual edges and semantic edges to describe the connection strength in terms of visual appearance and semantic content. Second, we evolve the graph by updating connection strength based on the predicted results of concept learning. Finally, we present a growth algorithm for the multi-edge concept graph to handle cross-dataset concept learning. Driven by the predictions, the multi-edge concept graph can dynamically evolve over time by adjusting the connection strength to adapt better to the observations. In addition, our approach can be considered a weakly-supervised learning algorithm since no labeled concepts are employed for learning. Experimental results demonstrate that evolution can significantly improve the learning of fundamental visual concepts by$\text{14.2}\%$,$\text{7.9}\%$and$\text{12.7}\%$in terms of F1-score for the MSRC, VOC2012 and MSCOCO datasets, respectively, and that the proposed EMCG approach largely outperforms the compared approaches.
Youtian Du, Guangxun Zhang, Zhongmin Cai, Chang Su 0002
IEEE Trans. Multim.4
2020 I'm All Eyes and Ears: Exploring Effective Locators for Privacy Awareness in IoT Scenarios
abstract
With the proliferation of IoT devices, there are growing concerns about being sensed or monitored by these devices unawares, especially in places perceived as private. We explore the design space of IoT locators to help people physically find nearby IoT devices. We first conducted a survey to understand people's willingness, current practices, and challenges in finding IoT devices. Our survey findings motivated us to design and implement low-cost locators (visual, auditory, and contextualized pictures) to help people find nearby devices. Through an iterative design process and two rounds of experiments, we found that these locators greatly reduced people's search time over a baseline of no locators. Many participants found the visual and auditory locators enjoyable. Some participants also appropriated the use of our system for other purposes, e.g., to learn about new IoT devices, instead of for privacy awareness.
Yunpeng Song, Yun Huang 0003, Zhongmin Cai, Jason I. Hong
CHI3
2020 Gender recognition using motion data from multiple smart devices
Youtian Du, Zhongmin Cai
Expert Syst. Appl.3
2019 Learning edge weights in file co-occurrence graphs for malware detection
Weixuan Mao, Zhongmin Cai, Bo Zeng 0001, Xiaohong Guan
Data Min. Knowl. Discov.2
2019 Normal and Easy: Account Sharing Practices in the Workplace
abstract
Work is being digitized across all sectors, and digital account sharing has become common in the workplace. In this paper, we conduct a qualitative and quantitative study of digital account sharing practices in the workplace. Across two surveys, we examine the sharing process at work, probing what accounts people share, how and why they share those accounts, and identifying the major challenges people face in sharing accounts. Our results demonstrate that account sharing in the modern workplace serves as a norm rather than a simple workaround; centralizing collaborative activity and reducing boundary management effort are key motivations for sharing. But people still struggle with a lack of activity accountability and awareness, conflicts over simultaneous access, difficulties controlling access, and collaborative password use. Our work provides insights into the current difficulties people face in workplace collaboration with online account sharing, as a result of inappropriate designs that still assume a single-user model for accounts. We highlight opportunities for CSCW and HCI researchers and designers to better support sharing by multiple people in a more usable and secure way.
Yunpeng Song, Cori Faklaris, Zhongmin Cai, Jason I. Hong, Laura A. Dabbish
Proc. ACM Hum. Comput. Interact.3
2018 From big data to knowledge: A spatio-temporal approach to malware detection
Weixuan Mao, Zhongmin Cai, Yuan Yang 0003, Xiaohong Shi, Xiaohong Guan
Comput. Secur.2
2018 Probabilistically Inferring Attack Ramifications Using Temporal Dependence Network
abstract
There is an increasing need of assessing and mitigating the effects of successful attacks. Uncovering malicious and contaminated objects in an attacked computing system is referred to as identification of attack ramifications. Previous methods identify the attack ramifications by directly tracking information flows (or dependences) from the intrusion root (i.e., the entry point of an attack). They face challenges such as undetermined intrusion root and dependence explosion. In this paper, we present a novel, light-weight method capable of identifying attack ramifications without the knowledge of intrusion root and less subject to dependency explosion. The method utilizes a probabilistic reasoning approach to fuse evidence derived from a subset of objects whose security states are known. It first splits the lifetime of an object into consecutive time slices (object-slices) to profile how the security state of this object changes over time. Then, a temporal dependence network (TDN) is constructed from system call traces to correlate object-slices according to information flows between them. Based on that, a Bayesian network (BN) model is built to characterize the uncertainties of infection propagations in the TDN. Finally, the method adopts loopy belief propagation on the BN model to infer the security state of an object. We evaluate the proposed method using a large data set of 389 attacks launched by the real-world malware samples including sophisticated ones such as Stuxnet. Extensive experiments demonstrate that our method is able to identify attack ramifications with a 97.47% precision at 97.21% recall without the knowledge of intrusion root.
Yuan Yang 0003, Zhongmin Cai, Chunyan Wang 0012, Junjie Zhang 0004
IEEE Trans. Inf. Forensics Secur.2
2017 Tracking user information using motion data through smartphones
abstract
In this paper, we use smartphone motion sensors and user basic activities as new source of information to detect and measure some dynamic information of user such as age group, footwear type and the floor surface types, which are important information in human sensing tasks and applications for healthcare, marketing, recommender system and etc. A random forest classifier was used to classify 262features extracted from smart phone accelero-meter and gyroscope sensors' data. Empirical evaluation was performed on a large-scale dataset containing 510 subjects and more than 111470 activities files. The results show that we can achieve accuracies of more than 85%, 92.5% and 95% for the detection of user's age, footwear types and floor surface type respectively.
Aghil Esmaeili Kelishomi, Zhongmin Cai, Mohammad Hossein Shayesteh
IJCB2
2017 Multi-touch Authentication Using Hand Geometry and Behavioral Information
abstract
In this paper we present a simple and reliable authentication method for mobile devices equipped with multi-touch screens such as smart phones, tablets and laptops. Users are authenticated by performing specially designed multi-touch gestures with one swipe on the touchscreen. During this process, both hand geometry and behavioral characteristics are recorded in the multi-touch traces and used for authentication. By combining both geometry information and behavioral characteristics, we overcome the problem of behavioral variability plaguing many behavior based authentication techniques - which often leads to less accurate authentication or poor user experience - while also ensuring the discernibility of different users with possibly similar handshapes. We evaluate the design of the proposed authentication method thoroughly using a large multi-touch dataset collected from 161 subjects with an elaborately designed procedure to capture behavior variability. The results demonstrate that the fusion of behavioral information with hand geometry features produces effective resistance to behavioral variability over time while at the same time retains discernibility. Our approach achieves EER of 5.84% with only 5 training samples and the performance is further improved to EER of 1.88% with enough training. Security analyses are also conducted to demonstrate that the proposed method is resilient against common smartphone authentication threats such as smudge attack, shoulder surfing attack and statistical attack. Finally, user acceptance of the method is illustrated via a usability study.
Yunpeng Song, Zhongmin Cai, Zhi-Li Zhang
IEEE Symposium on Security and Privacy2
2017 Security importance assessment for system objects and malware detection
Weixuan Mao, Zhongmin Cai, Don Towsley, Xiaohong Guan
Comput. Secur.2
2017 Improvement of affine iterative closest point algorithm for partial registration
abstract
In this study, partial registration problem with outliers and missing data in the affine case is discussed. To solve this problem, a novel objective function is proposed based on bidirectional distance and trimmed strategy, and then a new affine trimmed iterative closest point algorithm is given. First, when bidirectional distance measurement is applied, the ill‐posed partial registration problem in the affine case is prevented. Second, the overlapping percentage is solved by using trimmed strategy which uses as many correct overlapping points as possible. The authors’ method computes the affine transformation, correspondence and overlapping percentage automatically at each iterative step. In this way, it handles partially overlapping registration with outliers and missing data in the affine case well. Experimental results demonstrate that their method is more robust and precise than the state‐of‐the‐art algorithms. It also has good convergence and similar running time with traditional algorithms.
Zhongmin Cai, Shaoyi Du
IET Comput. Vis.2
2017 Spotting anomalous ratings for rating systems by analyzing target users and items
Zhihai Yang, Zhongmin Cai, Yuan Yang 0003
Neurocomputing2
2017 Detecting abnormal profiles in collaborative filtering recommender systems
Zhihai Yang, Zhongmin Cai
J. Intell. Inf. Syst.2
2016 Detecting Anomalous Ratings Using Matrix Factorization for Recommender Systems
Zhihai Yang, Zhongmin Cai
WAIM (2)2
2016 Estimating user behavior toward detecting anomalous ratings in rating systems
Zhihai Yang, Zhongmin Cai, Xiaohong Guan
Knowl. Based Syst.2
2016 Re-scale AdaBoost for attack detection in collaborative filtering recommender systems
Zhihai Yang, Lin Xu 0001, Zhongmin Cai, Zongben Xu
Knowl. Based Syst.3
2016 MouseIdentity: Modeling Mouse-Interaction Behavior for a User Verification System
abstract
Analysis of mouse-interaction behaviors for identifying individual computer users has experienced growing interest from information security and biometric researchers. This paper presents a simple and efficient user verification system by modeling mouse-interaction behavior, which is accurate and competent for future deployment. For each mouse-operation sample, holistic attributes of mouse trajectories are first analyzed using a power transformation method, to derive a schematic representation of behavior eigenspace. Then, a propagation-based segmentation method is developed to model the detailed dynamic process of mouse movements by performing adaptive behavior segmentation and then characterizing each obtained segment using fine-grained procedural motion metrics. Both schematic and procedural cues from mouse-interaction behaviors may be used independently for verification, being fused at the decision level using combination rules. Analyses are conducted using data from 106 subjects with 21 200 mouse-operation samples. The verification system achieves a 1.96% false-rejection rate and a 1.18% false-acceptance rate with a short verification time (about 6 s) and lightweight system overload. Additional experiments on the effect of sample length and subject pool further examine the applicability of our verification system. We also compare the proposed approach with the state-of-the-art approaches for the data collected. Our findings suggest that mouse-interaction behaviors can enhance traditional authentication systems.
Chao Shen 0001, Zhongmin Cai, Xiaohong Guan, Roy A. Maxion
IEEE Trans. Hum. Mach. Syst.2
2015 Identifying Intrusion Infections via Probabilistic Inference on Bayesian Network
Yuan Yang 0003, Zhongmin Cai, Weixuan Mao, Zhihai Yang
DIMVA2
2015 Probabilistic Inference on Integrity for Access Behavior Based Malware Detection
Weixuan Mao, Zhongmin Cai, Don Towsley, Xiaohong Guan
RAID2
2014 Centrality metrics of importance in access behaviors and malware detections
abstract
System objects play different roles in a computer system and exhibit different degrees of importance with respect to system security. Identifying importance metrics can help us to develop more effective and efficient security protection methods. However, there is little previous work on evaluating the importance of objects from the perspective of security. In this paper, we propose a novel approach to evaluate the importance of various system objects based on a bipartite dependency network representation of access behaviors observed in a computer system. We introduce centrality metrics from network science to quantitatively measure the relative importance of system objects and reveal their inherent connections to security properties such as integrity and confidentiality. Furthermore, we propose importance-metric based models to characterize process behaviors and identify abnormal access patterns with respect to confidentiality and integrity. Extensive experimental results on one real-world dataset demonstrate that our model is capable of detecting 7,257 malware samples from 27,840 benign processes at 93.94% TPR under 0.1% FPR. Moreover, a selective protection scheme based on a partial behavioral model of important objects achieves comparable or even better results in malware detection when compared with complete behavior models. This demonstrates the feasibility of the devised importance metrics and presents a promising new approach to malware detection.
Weixuan Mao, Zhongmin Cai, Xiaohong Guan, Don Towsley
ACSAC2
2014 Performance evaluation of anomaly-detection algorithms for mouse dynamics
Chao Shen 0001, Zhongmin Cai, Xiaohong Guan, Roy A. Maxion
Comput. Secur.2
2014 Mitigating Behavioral Variability for Mouse Dynamics: A Dimensionality-Reduction-Based Approach
abstract
Mouse dynamics is the process of identifying individual users on the basis of their mouse operating behaviors. Mouse dynamics analysis techniques do not provide an acceptable level of accuracy, perhaps due to behavioral variability. This study presents a dimensionality-reduction-based approach to mitigate the behavioral variability of mouse dynamics and improve the performance of mouse-dynamics-based continuous authentication. Variability was measured over the schematic features and motor-skill features extracted from each mouse behavior data session. A unified framework of employing dimensionality reduction methods (Multidimensional Scaling, Laplacian Eigenmap, Isometric Feature Mapping, and Local Linear Embedding) was developed to reduce behavioral variability by obtaining predominant characteristics from the original feature space. Classification techniques (Random Forest, Support Vector Machine, Neural Network, and Nearest Neighbor) were applied to the transformed feature space to perform the authentication task. Analyses were conducted using data from 840 half-hour sessions of 28 participants. Results indicated that for sufficiently long sequences, the transformed feature spaces had much less variability and the corresponding authentication performance was better than the original feature space with improvements of the false-acceptance rate by 89.6% and of the false-rejection rate by 77.4% in some cases. Additionally, an investigation of the relationships between variability and authentication error rates and detection time indicated that the variability and authentication error rates reduce greatly with the increase of detection time. For the data collected, the approach fared better than the state-of-the-art approaches. These findings suggest that variability reduction could improve mouse dynamics, so it may enhance current authentication mechanisms.
Zhongmin Cai, Chao Shen 0001, Xiaohong Guan
IEEE Trans. Hum. Mach. Syst.1
2013 Protect sensitive sites from phishing attacks using features extractable from inaccessible phishing URLs
abstract
Phishing is the third cyber-security threat globally and the first cyber-security threat in China. There were 61.69 million phishing victims in China alone from June 2011 to June 2012, with the total annual monetary loss more than 4.64 billion US dollars. These phishing attacks were highly concentrated in targeting at a few major Websites. Many phishing Webpages had a very short life span. In this paper, we assume the Websites to protect against phishing attacks are known, and study the effectiveness of machine learning based phishing detection using only lexical and domain features, which are available even when the phishing Webpages are inaccessible. We propose several novel highly effective features, and use the real phishing attack data against Taobao and Tencent, two main phishing targets in China, in studying the effectiveness of each feature, and each group of features. We then select an optimal set of features in our phishing detector, which has achieved a detection rate better than 98%, with a false positive rate of 0.64% or less. The detector is still effective when the distribution of phishing URLs changes.
Weibo Chu, Bin B. Zhu, Xiaohong Guan, Zhongmin Cai
ICC5
2013 On User Interaction Behavior as Evidence for Computer Forensic Analysis
Chao Shen 0001, Zhongmin Cai, Roy A. Maxion, Xiaohong Guan
IWDW2
2013 Real-time volume control for interactive network traffic replay
Weibo Chu, Xiaohong Guan, Zhongmin Cai, Lixin Gao 0001
Comput. Networks3
2013 Multi-view semi-supervised web image classification via co-graph
Youtian Du, Qian Li 0024, Zhongmin Cai, Xiaohong Guan
Neurocomputing3
2013 User Authentication Through Mouse Dynamics
abstract
Behavior-based user authentication with pointing devices, such as mice or touchpads, has been gaining attention. As an emerging behavioral biometric, mouse dynamics aims to address the authentication problem by verifying computer users on the basis of their mouse operating styles. This paper presents a simple and efficient user authentication approach based on a fixed mouse-operation task. For each sample of the mouse-operation task, both traditional holistic features and newly defined procedural features are extracted for accurate and fine-grained characterization of a user's unique mouse behavior. Distance-measurement and eigenspace-transformation techniques are applied to obtain feature components for efficiently representing the original mouse feature space. Then a one-class learning algorithm is employed in the distance-based feature eigenspace for the authentication task. The approach is evaluated on a dataset of 5550 mouse-operation samples from 37 subjects. Extensive experimental results are included to demonstrate the efficacy of the proposed approach, which achieves a false-acceptance rate of 8.74%, and a false-rejection rate of 7.69% with a corresponding authentication time of 11.8 seconds. Two additional experiments are provided to compare the current approach with other approaches in the literature. Our dataset is publicly available to facilitate future research.
Chao Shen 0001, Zhongmin Cai, Xiaohong Guan, Youtian Du, Roy A. Maxion
IEEE Trans. Inf. Forensics Secur.2
2012 Continuous authentication for mouse dynamics: A pattern-growth approach
abstract
Mouse dynamics is the process of identifying individual users based on their mouse operating characteristics. Although previous work has reported some promising results, mouse dynamics is still a newly emerging technique and has not reached an acceptable level of performance. One of the major reasons is intrinsic behavioral variability. This study presents a novel approach by using pattern-growth-based mining method to extract frequent-behavior segments in obtaining stable mouse characteristics, employing one-class classification algorithms to perform the task of continuous user authentication. Experimental results show that mouse characteristics extracted from frequent-behavior segments are much more stable than those from holistic behavior, and the approach achieves a practically useful level of performance with FAR of 0.37% and FRR of 1.12%. These findings suggest that mouse dynamics suffice to be a significant enhancement for a traditional authentication system. Our dataset is publicly available to facilitate future research.
Chao Shen 0001, Zhongmin Cai, Xiaohong Guan
DSN2
2012 Model-based real-time volume control for interactive network traffic replay
abstract
Traffic volume control is one of the fundamental requirements in traffic generation and transformation. However, due to the complex interactions between the generated traffic and replay environment (delay, packet loss, connection blocking, etc), controlling traffic volume in interactive network traffic replay becomes a challenging problem. In this paper, we present a novel model-based analytical method to address this problem where the generated traffic volume is regulated through adjustment of input traffic volume. By analyzing the replay mechanism in terms of how packets are processed, and properly choosing buffered packets amount and to-be-received packets amount as system states, we present a novel model-based analytical method to obtain the desired input volume. The traffic volume control problem is then converted to a state prediction problem where we employ Recursive Least Square (RLS) filter to predict system states. As compared to other adaptive control techniques, our method does not involve any learning scheme and hence completely requires no convergence time. Experimental studies further indicate that our method is efficient in tracking target traffic volume (both static and time-varying) and works under a wide range of network conditions.
Weibo Chu, Xiaohong Guan, Lixin Gao 0001, Zhongmin Cai
NOMS4
2011 Poster: can it be more practical?: improving mouse dynamics biometric performance
Chao Shen 0001, Zhongmin Cai, Xiaohong Guan
CCS2
2010 Balance Based Performance Enhancement for Interactive TCP Traffic Replay
abstract
Interactive network traffic replay plays an important role in testing and evaluating in-line network security devices such as Firewalls, IPSs, etc. In this paper we present a balance-based method for improving the performance of interactive TCP traffic replay. The new method is based on the inherent feature of the TCP protocol, that is, two communicating peers keep synchronized with each other using data acknowledgment. This feature is converted as a balance mechanism in interactive TCP traffic replay and incorporated into the current state-based method. In this way, the cost of state-checking can be significantly reduced and the replay performance is thus enhanced. To validate the effectiveness of the method we implement it by building an interactive replay system. The experimental results indicate that: 1) balance-checking reduces the overhead of state-checking by 40%; 2) the balance-based method enhances the overall replay performance by an average of 5% when the actual TCP traffic traces are replayed.
Weibo Chu, Xiaohong Guan, Zhongmin Cai, Mingxu Chen
ICC3
2010 Enhancing Web Page Classification via Local Co-training
abstract
In this paper we propose a new multi-view semi-supervised learning algorithm called Local Co-Training(LCT). The proposed algorithm employs a set of local models with vector outputs to model the relations among examples in a local region on each view, and iteratively refines the dominant local models (i.e. the local models related to the unlabeled examples chosen for enriching the training set) using unlabeled examples by the co-training process. Compared with previous co-training style algorithms, local co-training has two advantages: firstly, it has higher classification precision by introducing local learning; secondly, only the dominant local models need to be updated, which significantly decreases the computational load. Experiments on WebKB and Cora datasets demonstrate that LCT algorithm can effectively exploit unlabeled data to improve the performance of web page classification.
Youtian Du, Xiaohong Guan, Zhongmin Cai
ICPR3
2009 Feature Analysis of Mouse Dynamics in Identity Authentication and Monitoring
abstract
Mouse dynamics has recently become an interesting new topic in the area of behavioral biometrics due to its non-intrusiveness and convenience. Some promising results have been shown by previous researches on identity authentication and monitoring using characteristics in users' mouse actions. This paper explores mouse dynamics further by focusing on an important issue not addressed previously: behavioral variability. With an empirical study of long term behaviors of 10 computer users, we show variations are obvious in mouse activities and can have a serious impact if not considered carefully. To tackle the problem of variability, we propose a dimensionality reduction based approach which is demonstrated to be effective in our experiments. More specifically, the classification results after preprocessing by PCA and ISOMAP are shown to be much better than direct classification. Moreover, the results of a false acceptance rate (FAR) 0.55% and false rejection rate (FRR) 3.00% by the nonlinear method ISOMAP are comparable to the best result reported in literature while being subject to more behavioral variability.
Chao Shen 0001, Zhongmin Cai, Xiaohong Guan, Huilan Sha, Jingzi Du
ICC2
2003 A rough set theory based method for anomaly intrusion detection in computer network systems
abstract
Abstract: Intrusion detection is important in the defense‐in‐depth network security framework. This paper presents an effective method for anomaly intrusion detection with low overhead and high efficiency. The method is based on rough set theory to extract a set of detection rules with a minimal size as the normal behavior model from the system call sequences generated during the normal execution of a process. It is capable of detecting the abnormal operating status of a process and thus reporting a possible intrusion. Compared with other methods, the method requires a smaller size of training data set and less effort to collect training data and is more suitable for real‐time detection. Empirical results show that the method is promising in terms of detection accuracy, required training data set and efficiency.
Zhongmin Cai, Xiaohong Guan, Ping Shao, Qingke Peng, Guoji Sun
Expert Syst. J. Knowl. Eng.1