VLDB 2026 Research / reviewers in the wild / expert
Minjie Cai
dblp:164/8231
· DBLP profile ↗
20ranked-venue papers
5as first author
12since 2021 · last 2026
0000-0002-6688-3710ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-authorSecurity and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Tool-Assisted CVSS Vulnerability Scoring: A Controlled Quantitative Study of Human AssessmentabstractQuantitative vulnerability assessment is central to security management, guiding how risks are prioritized and mitigated. Yet, severity scoring relies on human judgment and is therefore subject to differences in experience, interpretation, and diligence; prior work has even shown expert disagreement. We examine an NLP-based assistive tool that visualizes keyword cues during assessment. In a controlled survey of 389 participants recruited via Amazon MTurk and Prolific, we statistically analyze how participant skills/demographics, vulnerability characteristics, and tool support affect outcomes. Results show the tool does not consistently improve assessment accuracy across expertise levels, but can help for specific vulnerability types (e.g., CWE-787) and CVSS metrics (AC, PR, Scope), and can increase user confidence. Beyond immediate performance, the tool can support training for manual assessment tasks that are hard to automate, as learning effects yield significant improvements on subsequent tasks. This work informs the design of cybersecurity decision-support tools and motivates future research on security training and human-centered security. Minjie Cai, Lianying Zhao, Xavier de Carné de Carnavalet, Fabio Massacci, Mengyuan Zhang 0001 |
CHI | 2 |
| 2025 | Egocentric Speaker Diarization with Vision-Guided Clustering and Adaptive Speech Re-detectionabstractSpeaker diarization aims to identify "who spoke when" in multi-person conversational scenarios. State-of-the-art audio-only diarization methods divide the task into multi-stages of speech segmentation, neural speaker embedding and unsupervised clustering. Egocentric speaker diarization (i.e., diarization in egocentric videos) is characterized by natural conversational scenarios involving a variety of noisy backgrounds, changing sound levels and overlapping speech. These cause difficulty in predicting the number of speakers from audio input, and heavily influence the performance of audio-based clustering. Although audio-visual modeling has been studied recently to enhance speaker diarization, unreliable visual information in egocentric videos may even degrade the performance. In this work, we propose a unified audio-visual diarization framework by incorporating visual guidance into the audio-only diarization pipeline. In addition, we also propose an adaptive speech re-detection strategy to detect and assign speaker identity to the mistakenly undetected audio segments. Experiments on the Ego4D dataset show that our method achieves state-of-the-art diarization performance in challenging egocentric scenarios. The code and model weights are available at https://github.com/YellowRiver2001/EgoDiarization. Daibo Liu, Minjie Cai |
ICASSP | 5 |
| 2025 | SiMHand: Mining Similar Hands for Large-Scale 3D Hand Pose Pre-trainingabstractWe present a framework for pre-training of 3D hand pose estimation from in-the-wild hand images sharing with similar hand characteristics, dubbed SiMHand. Pre-training with large-scale images achieves promising results in various tasks, but prior methods for 3D hand pose pre-training have not fully utilized the potential of diverse hand images accessible from in-the-wild videos. To facilitate scalable pre-training, we first prepare an extensive pool of hand images from in-the-wild videos and design our pre-training method with contrastive learning. Specifically, we collect over 2.0M hand images from recent human-centric videos, such as 100DOH and Ego4D. To extract discriminative information from these images, we focus on the similarity of hands: pairs of non-identical samples with similar hand poses. We then propose a novel contrastive learning method that embeds similar hand pairs closer in the feature space. Our method not only learns from similar samples but also adaptively weights the contrastive learning loss based on inter-sample distance, leading to additional performance gains. Our experiments demonstrate that our method outperforms conventional contrastive learning approaches that produce positive pairs solely from a single image with data augmentation. We achieve significant improvements over the state-of-the-art method (PeCLR) in various datasets, with gains of 15% on FreiHand, 10% on DexYCB, and 4% on AssemblyHands. Our code is available at https://github.com/ut-vision/SiMHand. Nie Lin, Takehiko Ohkawa, Yifei Huang 0002, Mingfang Zhang 0002, Minjie Cai, Ryosuke Furuta, Yoichi Sato 0001 |
ICLR | 5 |
| 2025 | DFL: A DOM sample generation oriented fuzzing framework for browser rendering engines
Guoyun Duan, Minjie Cai, Jianhua Sun 0002, Hao Chen 0002 |
Inf. Softw. Technol. | 3 |
| 2024 | MaDroid: A maliciousness-aware multifeatured dataset for detecting android malware
Guoyun Duan, Minjie Cai, Jianhua Sun 0002, Hao Chen 0002 |
Comput. Secur. | 3 |
| 2024 | Uncertainty-Aware and Class-Balanced Domain Adaptation for Object Detection in Driving ScenesabstractThis work tackles the cross-domain object detection problem which aims to generalize a pre-trained object detector to different domains (driving scenes) without labels. An uncertainty-aware and class-balanced domain adaptation method is proposed based on two motivations: 1) estimation and exploitation of model uncertainty in a new domain is critical for reliable domain adaptation; and 2) in domain adaptation the distribution alignment of two domains as well as the maintaining of category discriminability are both important. In particular, we compose a Bayesian CNN-based framework for uncertainty estimation in object detection. We propose an algorithm for generating uncertainty-aware pseudo-labels, which are then used in uncertainty-guided self-training and category-aware feature alignment. We further devise a scheme with class-balanced memory banks to address the long-tail distribution problem in category-aware feature alignment. Experiments on multiple cross-domain object detection benchmarks show that our proposed method achieves state-of-the-art performance. Minjie Cai, Jianaresi Kezierbieke, Xionghu Zhong, Hao Chen 0002 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2023 | The Flaw Within: Identifying CVSS Score Discrepancies in the NVDabstractCloud security frameworks, like OpenSCAP, rely on vulnerability databases such as the National Vulnerability Database (NVD) to assess threats, ensure compliance, and manage patches efficiently. However, despite their popularity, vulnerability databases are not exempt from errors. Prior research showed inconsistencies between multiple databases, as well as incorrect software or vendor names, and publication dates. In this study, we discovered and proposed a systematic approach to detect a new form of inconsistency whereby entries with identical or semantically similar vulnerability descriptions are assigned distanced scores, which can skew risk assessments, and potentially misguide mitigation strategies. Our analysis identified 12,866 entries suffering from such inconsistencies, highlighting the most error-prone Common Vulnerability Scoring System (CVSS) metrics and vulnerability types, as well as the observed score deviation. We believe our study can bring this inconsistency issue to the community’s attention and pave the way for further investigation thereof. Minjie Cai, Mengyuan Zhang 0001, Lianying Zhao, Xavier de Carné de Carnavalet |
CloudCom | 2 |
| 2023 | Blind Estimation of Room Impulse Response from Monaural Reverberant Speech with Segmental Generative Neural Network
Zhiheng Liao, Feifei Xiong, Juan Luo, Minjie Cai, Chng Eng Siong, Jinwei Feng, Xionghu Zhong |
INTERSPEECH | 4 |
| 2023 | Emotion-Aware Audio-Driven Face Animation via Contrastive Feature Disentanglement
Juan Luo, Xionghu Zhong, Minjie Cai |
INTERSPEECH | 4 |
| 2023 | DongTing: A large-scale dataset for anomaly detection of the Linux kernel
Guoyun Duan, Yuanzhi Fu, Minjie Cai, Hao Chen 0002, Jianhua Sun 0002 |
J. Syst. Softw. | 3 |
| 2023 | First- And Third-Person Video Co-Analysis By Learning Spatial-Temporal Joint AttentionabstractRecent years have witnessed a tremendous increase of first-person videos captured by wearable devices. Such videos record information from different perspectives than the traditional third-person view, and thus show a wide range of potential usages. However, techniques for analyzing videos from different views can be fundamentally different, not to mention co-analyzing on both views to explore the shared information. In this paper, we take the challenge of cross-view video co-analysis and deliver a novel learning-based method. At the core of our method is the notion of "joint attention", indicating the shared attention regions that link the corresponding views, and eventually guide the shared representation learning across views. To this end, we propose a multi-branch deep network, which extracts cross-view joint attention and shared representation from static frames with spatial constraints, in a self-supervised and simultaneous manner. In addition, by incorporating the temporal transition model of the joint attention, we obtain spatial-temporal joint attention that can robustly capture the essential information extending through time. Our method outperforms the state-of-the-art on the standard cross-view video matching tasks on public datasets. Furthermore, we demonstrate how the learnt joint information can benefit various applications through a set of qualitative and quantitative experiments. Huangyue Yu, Minjie Cai, Yunfei Liu 0001, Feng Lu 0005 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Audio-Visual Event Localization by Learning Spatial and Semantic Co-AttentionabstractThis work aims to temporally localize events that are both audible and visible in video. Previous methods mainly focused on temporal modeling of events with simple fusion of audio and visual features. In natural scenes, a video records not only the events of interest but also ambient acoustic noise and visual background, resulting in redundant information in the raw audio and visual features. Thus, direct fusion of the two features often causes false localization of the events. In this paper, we propose a co-attention model to exploit the spatial and semantic correlations between the audio and visual features, which helps guide the extraction of discriminative features for better event localization. Our assumption is that in an audio-visual event, shared semantic information between audio and visual features exists and can be extracted by attention learning. Specifically, the proposed co-attention model is composed of a co-spatial attention module and a co-semantic attention module that are used to model the spatial and semantic correlations, respectively. The proposed co-attention model can be applied to various event localization tasks, such as cross-modality localization and multimodal event localization. Experiments on the public audio-visual event (AVE) dataset demonstrate that the proposed method achieves state-of-the-art performance by learning spatial and semantic co-attention. Xionghu Zhong, Minjie Cai, Hao Chen 0002, Wenwu Wang 0001 |
IEEE Trans. Multim. | 3 |
| 2020 | Generalizing Hand Segmentation in Egocentric Videos With Uncertainty-Guided Model AdaptationabstractAlthough the performance of hand segmentation in egocentric videos has been significantly improved by using CNNs, it still remains a challenging issue to generalize the trained models to new domains, e.g., unseen environments. In this work, we solve the hand segmentation generalization problem without requiring segmentation labels in the target domain. To this end, we propose a Bayesian CNN-based model adaptation framework for hand segmentation, which introduces and considers two key factors: 1) prediction uncertainty when the model is applied in a new domain and 2) common information about hand shapes shared across domains. Consequently, we propose an iterative self-training method for hand segmentation in the new domain, which is guided by the model uncertainty estimated by a Bayesian CNN. We further use an adversarial component in our framework to utilize shared information about hand shapes to constrain the model adaptation process. Experiments on multiple egocentric datasets show that the proposed method significantly improves the generalization performance of hand segmentation. Minjie Cai, Feng Lu 0005, Yoichi Sato 0001 |
CVPR | 1 |
| 2020 | An Ego-Vision System for Discovering Human Joint AttentionabstractJoint attention often happens during social interactions, in which individuals share focus on the same object. This article proposes an egocentric vision-based system (ego-vision system) that aims to discover the objects looked at jointly by a group of persons engaged in interactive activities. The proposed system relies on a collection of wearable eye-tracking cameras that provide an egocentric view of the interaction scenes as well as points-of-gaze measurement of each participant. Technically in our system, we develop a hierarchical conditional random field (CRF) based graphical model that can temporally localize joint attention periods and spatially segment objects of joint attention. By solving these two coupled tasks together in an iterative optimization procedure, we show that human joint attention can be reliably discovered from videos even with cluttered background and noisy gaze measurement. A new dataset of joint attention is collected and annotated for evaluating the two tasks of joint attention where two to four persons are involved. Experimental results demonstrate that our approach achieves state-of-the-art performance on both tasks of spatial segmentation and temporal localization of joint attention. Yifei Huang 0002, Minjie Cai, Yoichi Sato 0001 |
IEEE Trans. Hum. Mach. Syst. | 2 |
| 2020 | Mutual Context Network for Jointly Estimating Egocentric Gaze and ActionabstractIn this work, we address two coupled tasks of gaze prediction and action recognition in egocentric videos by exploring their mutual context: the information from gaze prediction facilitates action recognition and vice versa. Our assumption is that during the procedure of performing a manipulation task, on the one hand, what a person is doing determines where the person is looking at. On the other hand, the gaze location reveals gaze regions which contain important and information about the undergoing action and also the non-gaze regions that include complimentary clues for differentiating some fine-grained actions. We propose a novel mutual context network (MCN) that jointly learns action-dependent gaze prediction and gaze-guided action recognition in an end-to-end manner. Experiments on multiple egocentric video datasets demonstrate that our MCN achieves state-of-the-art performance of both gaze prediction and action recognition. The experiments also show that action-dependent gaze patterns could be learned with our method. Yifei Huang 0002, Minjie Cai, Zhenqiang Li 0002, Feng Lu 0005, Yoichi Sato 0001 |
IEEE Trans. Image Process. | 2 |
| 2019 | What I See Is What You See: Joint Attention Learning for First and Third Person Video Co-analysisabstractIn recent years, more and more videos are captured from the first-person viewpoint by wearable cameras. Such first-person video provides additional information besides the traditional third-person video, and thus has a wide range of applications. However, techniques for analyzing the first-person video can be fundamentally different from those for the third-person video, and it is even more difficult to explore the shared information from both viewpoints. In this paper, we propose a novel method for first- and third-person video co-analysis. At the core of our method is the notion of "joint attention'', indicating the learnable representation that corresponds to the shared attention regions in different viewpoints and thus links the two viewpoints. To this end, we develop a multi-branch deep network with a triplet loss to extract the joint attention from the first- and third-person videos via self-supervised learning. We evaluate our method on the public dataset with cross-viewpoint video matching tasks. Our method outperforms the state-of-the-art both qualitatively and quantitatively. We also demonstrate how the learned joint attention can benefit various applications through a set of additional experiments. Huangyue Yu, Minjie Cai, Yunfei Liu 0001, Feng Lu 0005 |
ACM Multimedia | 2 |
| 2019 | Desktop Action Recognition From First-Person Point-of-ViewabstractDesktop action recognition from first-person view (egocentric) video is an important task due to its omnipresence in our daily life, and the ideal first-person viewing perspective for observing hand-object interactions. However, no previous research efforts have been dedicated on the benchmark of the task. In this paper, we first release a dataset of daily desktop actions recorded with a wearable camera and publish it as a benchmark for desktop action recognition. Regular desktop activities of six participants were recorded in egocentric video with a wide-angle head-mounted camera. In particular, we focus on five common desktop actions in which hands are involved. We provide original video data, action annotations at frame-level, and hand masks at pixel-level. We also propose a feature representation for the characterization of different desktop actions based on the spatial and temporal information of hands. In experiments, we illustrate the statistical information about the dataset, and evaluate the action recognition performance of different features as a baseline. The proposed method achieves promising performance for five action classes. Minjie Cai, Feng Lu 0005, Yue Gao 0002 |
IEEE Trans. Cybern. | 1 |
| 2018 | Predicting Gaze in Egocentric Video by Learning Task-Dependent Attention Transition
Yifei Huang 0002, Minjie Cai, Zhenqiang Li 0002, Yoichi Sato 0001 |
ECCV (4) | 2 |
| 2017 | An Ego-Vision System for Hand Grasp AnalysisabstractThis paper presents an egocentric vision (ego-vision) system for hand grasp analysis in unstructured environments. Our goal is to automatically recognize hand grasp types and to discover the visual structures of hand grasps using a wearable camera. In the proposed system, free hand–object interactions are recorded from a first-person viewing perspective. State-of-the-art computer vision techniques are used to detect hands and extract hand-based features. A new feature representation that incorporates hand tracking information is also proposed. Then, grasp classifiers are trained to discriminate among different grasp types from a predefined grasp taxonomy. Based on the trained grasp classifiers, visual structures of hand grasps are learned using an iterative grasp clustering method. In experiments, grasp recognition performance in both laboratory and real-world scenarios is evaluated. The best classification accuracy our system achieves is $\text{92}\%$ and $\text{59}\%$ , respectively. System generality to different tasks and users is also verified by the experiments. Analysis in a real-world scenario shows that it is possible to automatically learn intuitive visual grasp structures that are consistent with expert-designed grasp taxonomies. Minjie Cai, Kris Makoto Kitani, Yoichi Sato 0001 |
IEEE Trans. Hum. Mach. Syst. | 1 |
| 2015 | A scalable approach for understanding the visual structures of hand graspsabstractOur goal is to automatically recognize hand grasps and to discover the visual structures (relationships) between hand grasps using wearable cameras. Wearable cameras provide a first-person perspective which enables continuous visual hand grasp analysis of everyday activities. In contrast to previous work focused on manual analysis of first-person videos of hand grasps, we propose a fully automatic vision-based approach for grasp analysis. A set of grasp classifiers are trained for discriminating between different grasp types based on large margin visual predictors. Building on the output of these grasp classifiers, visual structures among hand grasps are learned based on an iterative discriminative clustering procedure. We first evaluated our classifiers on a controlled indoor grasp dataset and then validated the analytic power of our approach on real-world data taken from a machinist. The average F1 score of our grasp classifiers achieves over 0.80 for the indoor grasp dataset. Analysis of real-world video shows that it is possible to automatically learn intuitive visual grasp structures that are consistent with expert-designed grasp taxonomies. Minjie Cai, Kris Makoto Kitani, Yoichi Sato 0001 |
ICRA | 1 |