Zhenyu Chen 0003

dblp:86/541-3 · DBLP profile ↗
← Back
30ranked-venue papers
1as first author
11since 2021 · last 2026
0000-0002-4989-7109ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 9 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 2 since 2021Computer networks · 5 · 2 since 2021Human-computer interaction and ubiquitous computing · 3Databases, data management, data science and information retrieval · 2 · 1 first-author
YearPublicationVenuePosition
2026 Prompt-Free Efficient Adaptation of Segment Anything Model for Remote Sensing Landslide Detection
abstract
Landslides pose serious risks to human safety and infrastructure, making accurate and timely detection essential for disaster assessment. Although deep learning has advanced landslide recognition from remote sensing imagery, existing methods typically rely on task-specific designs, which limit their generalization ability across diverse geographic environments. The Segment Anything Model (SAM) offers strong generalization and zero-shot segmentation ability, yet its dependence on manually provided prompts and limited suitability for remote sensing imagery constrain its effectiveness in landslide detection. To address these limitations, we propose PF-SAM, a prompt-free and parameter-efficient adaptation framework tailored for landslide segmentation. PF-SAM eliminates manual prompts by autonomously constructing the prompt embeddings that drive SAM’s segmentation process and adapting SAM to landslide remote sensing data. Specifically, we first propose the CNN-Token Adapter, which extracts multi-scale CNN features and transforms them into prompt tokens, providing SAM with localized geometric and textural cues essential for detecting landslides of diverse shapes and sizes. Then, we propose the Global Prompt Generation Mechanism (GPGM) that further enhances segmentation: the Dense Prompt Generation Module (DPGM) generates the coarse mask as the dense prompt to guide SAM toward the correct landslide regions, while the Sparse Prompt Generation Module (SPGM) selects high-confidence embeddings to generate sparse prompts that enhance the precise localization of landslide regions. Extensive experiments on multiple landslide datasets show that PF-SAM achieves state-of-the-art performance, substantially surpassing existing methods in detection accuracy and robustness.
Zuolei Li, Xingyu Gao 0001, Zhenyu Chen 0003, Yawen Duan
IEEE Trans. Circuits Syst. Video Technol.3
2025 Deep Learning to Hash With Application to Cross-View Nearest Neighbor Search
abstract
Learning hash functions for approximate nearest neighbor search of high-dimensional data has received a surge of interests in recent years. Most existing methods are often concerned with learning hash functions for nearest neighbor search on high-dimensional data from a single source. In many real-world applications, data can be collected from diverse sources or represented using different feature descriptors. This raises an open challenge, i.e., the Cross-View Nearest Neighbor Search (CVNNS), where the representation of a query instance can be different from that of target instances to be retrieved in database. The key challenge of cross-view search is to learn an effective shared representation which can effectively connect the query instance and the target instances to be retrieved. In this paper, we present a new cross-view nearest neighbor search scheme by applying the emerging deep learning to hash techniques. In particular, we investigate two different architectures of deep Restricted Boltzmann Machines (RBMs) for learning to hash toward cross-view nearest neighbor search, and conduct extensive experiments to examine their empirical performance on diverse settings of cross-view image retrieval tasks. The encouraging results show that our technique outperforms the state-of-the-art approaches.
Xingyu Gao 0001, Zhenyu Chen 0003, Boshen Zhang, Jianze Wei
IEEE Trans. Circuits Syst. Video Technol.2
2025 Scribble-Supervised Video Object Segmentation via Scribble Enhancement
abstract
Current video object segmentation methods heavily rely on pixel-level mask annotations when training, which are expensive and time-consuming to acquire. To address this problem, some approaches try to train with sparse scribble annotations and take sparse target scribble as initial information for inference. However, due to the sparsity of scribble annotations, the performance is often limited, and the corresponding loss function needs to be designed. Inspired by the powerful ability of Segment Anything Model (SAM) to leverage prompt for segmentation, we argue that this problem can be alleviated by improving the quality of scribble. Therefore, we propose SEVOS, a framework for scribble-supervised video object segmentation, which contains a scribble enhancement algorithm and an semi-supervised video object segmentation network. Specifically, the scribble enhancement algorithm first samples corresponding positive sample points and negative sample points from target scribbles, and then feeds them into the SAM in turn, achieving high-quality scribble enhancement without human intervention. This algorithm augments the scribble-annotated video dataset, which is used for additional training of the model. Furthermore, we design a post-processing enhancement algorithm to further improve the prediction results. The obtained model outperforms state-of-the-art methods with a considerable performance gap, indicating the generalization and effectiveness of the proposed model.
Xingyu Gao 0001, Zuolei Li, Hailong Shi, Zhenyu Chen 0003, Peilin Zhao
IEEE Trans. Circuits Syst. Video Technol.4
2025 Deep Mutual Distillation for Unsupervised Domain Adaptation Person Re-Identification
abstract
Unsupervised domain adaptation person re-identification (UDA person re-ID) aims at transferring the knowledge on the source domain with expensive manual annotation to the unlabeled target domain. Most of the recent papers leverage pseudo-labels for the target images to accomplish this task. However, the noise in the generated labels hinders the identification system from learning discriminative features. To address this problem, we propose a deep mutual distillation (DMD) to generate reliable pseudo-labels for UDA person re-ID. The proposed DMD applies two parallel branches for feature extraction, and each branch serves as the teacher of the other to generate pseudo-labels for its training. This mutually reinforcing optimization framework enhances the reliability of pseudo-labels, improving the identification performance. In addition, we present a bilateral graph representation (BGR) to describe the pedestrian images. BGR mimics the person re-identification of the human to aggregate the identity features according to the visual similarity and attribute consistency. Experimental results on Market-1501 and Duke demonstrate the effectiveness and generalization of the proposed method.
Xingyu Gao 0001, Zhenyu Chen 0003, Jianze Wei, Rubo Wang, Zhijun Zhao
IEEE Trans. Multim.2
2024 Knowledge Enhanced Vision and Language Model for Multi-Modal Fake News Detection
abstract
The rapid dissemination of fake news and rumors through the Internet and social media platforms poses significant challenges and raises concerns in the public sphere. Automatic detection of fake news plays a crucial role in mitigating the spread of misinformation. While recent approaches have focused on leveraging neural networks to improve textual and visual representations in multi-modal fake news analysis, they often overlook the potential of incorporating knowledge information to verify facts within news articles. In this paper, we propose a knowledge enhanced vision and language model for multi-modal fake news detection. Our proposed model integrates information from large scale open knowledge graphs to augment its ability to discern the veracity of news content. Unlike previous methods that utilize separate models to extract textual and visual features, we synthesize a unified model capable of extracting both types of features simultaneously. To represent news articles, we introduce a graph structure where nodes encompass entities, relationships extracted from the textual content, and objects depicted in associated images. By utilizing the knowledge graph, we establish meaningful relationships between nodes within the news articles. Experimental evaluations on a real-world multi-modal dataset from Twitter demonstrate significant performance improvement by incorporating knowledge information.
Xingyu Gao 0001, Xi Wang 0014, Zhenyu Chen 0003, Wei Zhou 0019, Steven C. H. Hoi
IEEE Trans. Multim.3
2024 Self-Adaptive Graph With Nonlocal Attention Network for Skeleton-Based Action Recognition
abstract
Graph convolutional networks (GCNs) have achieved encouraging progress in modeling human body skeletons as spatial-temporal graphs. However, existing methods still suffer from two inherent drawbacks. Firstly, these models process the input data based on the physical structure of the human body, which leads to some latent correlations among joints being ignored. Furthermore, the key temporal relationships between nonadjacent frames are overlooked, preventing to fully learn the changes of the body joints along the temporal dimension. To address these issues, we propose an innovative spatial-temporal model by introducing a self-adaptive GCN (SAGCN) with global attention network, collectively termed SAGGAN. Specifically, the SAGCN module is proposed to construct two additional dynamic topological graphs to learn the common characteristics of all data and represent a unique pattern for each sample, respectively. Meanwhile, the global attention module (spatial attention (SA) and temporal attention (TA) modules) is designed to extract the global connections between different joints in a single frame and model temporal relationships between adjacent and nonadjacent frames in temporal sequences. In this manner, our network can capture richer features of actions for accurate action recognition and overcome the defect of the standard graph convolution. Extensive experiments on three benchmark datasets (NTU-60, NTU-120, and Kinetics) have demonstrated the superiority of our proposed method.
Chen Pang 0001, Xingyu Gao 0001, Zhenyu Chen 0003, Lei Lyu 0001
IEEE Trans. Neural Networks Learn. Syst.3
2023 Contrastive Multi-Level Graph Neural Networks for Session-Based Recommendation
abstract
Session-based recommendation (SBR) aims to predict the next item at a certain time point based on anonymous user behavior sequences. Existing methods typically model session representation based on simple item transition information. However, since session-based data consists of limited users' short-term interactions, modeling session representation by capturing fixed item transition information from a single dimension suffers from data sparsity. In this paper, we propose a novel contrastive multi-level graph neural networks (CM-GNN) to better exploit complex and high-order item transition information. Specifically, CM-GNN applies local-level graph convolutional network (L-GCN) and global-level graph convolutional network (G-GCN) on the current session and all the sessions respectively, to effectively capture pairwise relations over all the sessions by aggregation strategy. Meanwhile, CM-GNN applies hyper-level graph convolutional network (H-GCN) to capture high-order information among all the item transitions. CM-GNN further introduces an attention-based fusion module to learn pairwise relation-based session representation by fusing the item representations generated by L-GCN and G-GCN. CM-GNN averages the item representations obtained by H-GCN to obtain high-order relation-based session representation. Moreover, to convert the high-order item transition information into the pairwise relation-based session representation, CM-GNN maximizes the mutual information between the representations derived from the fusion module and the average pool layer by contrastive learning paradigm. We conduct extensive experiments on several widely used benchmark datasets to validate the efficacy of the proposed method. The encouraging results demonstrate that our proposed method outperforms the state-of-the-art SBR techniques.
Fuyun Wang, Xingyu Gao 0001, Zhenyu Chen 0003, Lei Lyu 0001
IEEE Trans. Multim.3
2023 Dilated Convolution-based Feature Refinement Network for Crowd Localization
abstract
As an emerging computer vision task, crowd localization has received increasing attention due to its ability to produce more accurate spatially predictions. However, continuous scale variations in complex crowd scenes lead to tiny individuals at the edges, so that existing methods cannot achieve precise crowd localization. Aiming at alleviating the above problems, we propose a novel Dilated Convolution-based Feature Refinement Network (DFRNet) to enhance the representation learning capability. Specifically, the DFRNet is built with three branches that can capture the information of each individual in crowd scenes more precisely. More specifically, we introduce a Feature Perception Module to model long-range contextual information at different scales by adopting multiple dilated convolutions, thus providing sufficient feature information to perceive tiny individuals at the edge of images. Afterwards, a Feature Refinement Module is deployed at multiple stages of the three branches to facilitate the mutual refinement of feature information at different scales, thus further improving the expression capability of multi-scale contextual information. By incorporating the above modules, DFRNet can locate individuals in complex scenes more precisely. Extensive experiments on multiple datasets demonstrate that the proposed method has more advanced performance compared to existing methods and can be more accurately adapted to complex crowd scenes.
Xingyu Gao 0001, Jinyang Xie, Zhenyu Chen 0003, Anan Liu, Zhenan Sun, Lei Lyu 0001
ACM Trans. Multim. Comput. Commun. Appl.3
2023 Learning Semantic Representation on Visual Attribute Graph for Person Re-identification and Beyond
abstract
Person re-identification (re-ID) aims to match pedestrian pairs captured from different cameras. Recently, various attribute-based models have been proposed to combine the pedestrian attribute as an auxiliary semantic information to learn a more discriminative pedestrian representation. However, these methods usually directly concatenate the visual branch and attribute branch embeddings as the final pedestrian representation, which ignores the semantic relation between the pedestrian revealed by attribute similarity. To capture and explore such semantic relation, we propose a unified pedestrian representation framework, called Visual Attribute Graph Embedding Network (VAGEN), to simultaneously learn attribute and visual representation. We unify the visual embedding and attribute similarity into a Visual Attribute Graph, where pedestrian is considered as a node and attribute similarity as an edge. Then, we learn graph node embedding to generate pedestrian representation through Graph Neural Network. Except for this unified representation for visual and attribute embeddings, VAGEN also conducts implicitly hard example mining for visual similar false-positive results, which has not been explored yet among existing attribute-based methods. We conduct extensive empirical studies on several person re-ID datasets to evaluate our proposed algorithm from different aspects. The results show that our proposed method outperforms state-of-the-art techniques with considerable margins.
Geyu Tang, Xingyu Gao 0001, Zhenyu Chen 0003
ACM Trans. Multim. Comput. Commun. Appl.3
2022 Task-Adaptive Attention for Image Captioning
abstract
Attention mechanisms are now widely used in image captioning models. However, most attention models only focus on visual features. When generating syntax related words, little visual information is needed. In this case, these attention models could mislead the word generation. In this paper, we propose Task-Adaptive Attention module for image captioning, which can alleviate this misleading problem and learn implicit non-visual clues which can be helpful for the generation of non-visual words. We further introduce a diversity regularization to enhance the expression ability of the Task-Adaptive Attention module. Extensive experiments on the MSCOCO captioning dataset demonstrate that by plugging our Task-Adaptive Attention module into a vanilla Transformer-based image captioning model, performance improvement can be achieved.
Chenggang Yan 0001, Yiming Hao, Liang Li 0003, Jian Yin 0003, Anan Liu, Zhendong Mao 0001, Zhenyu Chen 0003, Xingyu Gao 0001
IEEE Trans. Circuits Syst. Video Technol.7
2021 Unsupervised adversarial domain adaptation with similarity diffusion for person re-identification
Geyu Tang, Xingyu Gao 0001, Zhenyu Chen 0003, Huicai Zhong
Neurocomputing3
2020 Learning salient features to prevent model drift for correlation tracking
Yu Zhang 0102, Xingyu Gao 0001, Zhenyu Chen 0003, Huicai Zhong, Liang Li 0003, Chenggang Yan 0001, Tao Shen 0004
Neurocomputing3
2020 Advancing Image Understanding in Poor Visibility Environments: A Collective Benchmark Study
abstract
Existing enhancement methods are empirically expected to help the high-level end computer vision task: however, that is observed to not always be the case in practice. We focus on object or face detection in poor visibility enhancements caused by bad weathers (haze, rain) and low light conditions. To provide a more thorough examination and fair comparison, we introduce three benchmark sets collected in real-world hazy, rainy, and low-light conditions, respectively, with annotated objects/faces. We launched the UG2+ challenge Track 2 competition in IEEE CVPR 2019, aiming to evoke a comprehensive discussion and exploration about whether and how low-level vision techniques can benefit the high-level automatic visual recognition in various scenarios. To our best knowledge, this is the first and currently largest effort of its kind. Baseline results by cascading existing enhancement and detection models are reported, indicating the highly challenging nature of our new data as well as the large room for further technical innovations. Thanks to a large participation from the research community, we are able to analyze representative team solutions, striving to better identify the strengths and limitations of existing mindsets as well as the future directions.
Wenhan Yang, Ye Yuan 0012, Wenqi Ren, Jiaying Liu 0001, Walter J. Scheirer, Zhangyang Wang, Taiheng Zhang, Qiaoyong Zhong, Di Xie, Shiliang Pu, Yuqiang Zheng, Yanyun Qu, Yuhong Xie, Hao Jiang 0014, Siyuan Yang 0001, Yan Liu 0041, Xiaochao Qu, Pengfei Wan 0001, Shuai Zheng 0005, Minhui Zhong, Taiyi Su, Lingzhi He, Yandong Guo, Yao Zhao 0001, Zhenfeng Zhu, Jinxiu Liang, Jingwen Wang 0003, Yuhui Quan, Yong Xu 0007, Bo Liu 0112, Xin Liu 0012, Tingyu Lin 0003, Xiaochuan Li 0001, Feng Lu 0005, Lin Gu 0003, Shengdi Zhou, Cong Cao 0005, Cheng Chi 0003, Chubin Zhuang, Zhen Lei 0001, Stan Z. Li, Shizheng Wang, Ruizhe Liu, Dong Yi, Zheming Zuo, Jianning Chi, Huan Wang 0014, Kai Wang 0036, Yixiu Liu, Xingyu Gao 0001, Zhenyu Chen 0003, Yongzhou Li, Huicai Zhong, Jing Huang 0017, Heng Guo 0003, Jianfei Yang 0001, Wenjuan Liao, Jiangang Yang, Liguo Zhou, Mingyue Feng, Likun Qin
IEEE Trans. Image Process.57
2020 Mining Spatial-Temporal Similarity for Visual Tracking
abstract
Correlation filter (CF) is a critical technique to improve accuracy and speed in the field of visual object tracking. Despite being studied extensively, most existing CF methods suffer from failing to make the most of the inherent spatial-temporal prior of videos. To address this limitation, as consecutive frames are eminently resemble in most videos, we investigate a novel scheme to predict targets' future state by exploiting previous observations. Specifically, in this paper, we propose a prediction based CF tracking framework by learning the spatial-temporal similarity of consecutive frames for sample managing, template regularization, and training response pre-weighting. We model the learning problem theoretically as a novel objective and provide effective optimization algorithms to solve the learning task. In addition, we implement two CF trackers with different features. Extensive experiments are conducted on three popular benchmarks to validate our scheme. The encouraging results demonstrate that the proposed scheme can significantly boost the accuracy of CF tracking, and the two trackers achieve competitive performances against state-of-the-art trackers. We finally present a comprehensive analysis on the efficacy of our proposed method and the efficiency of our trackers to facilitate real-world visual tracking applications.
Yu Zhang 0102, Xingyu Gao 0001, Zhenyu Chen 0003, Huicai Zhong, Hongtao Xie 0001, Chenggang Yan 0001
IEEE Trans. Image Process.3
2019 Learning Correlation Filter With Detection Response For Visual Tracking
abstract
Correlation Filter (CF) has been a powerful tool for real-time visual object tracking. Although enormous improvements have been made since the first CF tracker was proposed, most of the existing CF tracking schemes adopt a fixed Gaussian label to train correlation filter. We argue that it is not optimal for challenging scenarios. Additionally, existing training label managing strategies are too complex to maintain the speed advantage of CF algorithms. In this work, we propose a novel label to supervise filter training. Our method is capable of being integrated into most CF trackers conveniently without increasing computational complexity. We employ the proposed scheme for four existing CF trackers and conduct extensive experiments on popular benchmark. The encouraging experimental results validate both the effi-ciency and efficacy of our proposed method. What's more, we provide a detailed analysis for challenging applications and hyperparameter setting.
Yu Zhang 0102, Xingyu Gao 0001, Zhenyu Chen 0003, Huicai Zhong
ICIP3
2019 Multimodel Framework for Indoor Localization Under Mobile Edge Computing Environment
abstract
Location estimation technology under the wireless environment has become a vital technology in the field of mobile edge computing. Especially, under the mobile edge of entire networks environment, indoor location estimation is gradually getting the interest research and application topic, due to technical constraints of global positioning system technology for indoor environment and the popularity of the mobile edge computing servers. In this paper, the widely used single-model framework for indoor localization is presented as an introduction, which consists of three stages: 1) sample data collection; 2) model building; and 3) localization estimation. And then, through analyzing of the actual scene of indoor localization, a new framework for indoor localization under mobile edge computing environment, named Multimodel, is proposed from the theoretical perspective. It is mainly based on the observation that the environment of the sample data collection and that of localization data collection may change seriously. In order to make up for the shortcomings of this framework, two combinatorial optimization problems are proposed. Later, we discuss the NP-hardness of them in several different cases. In addition, two heuristic algorithms are given, and the performance of which are illustrated by the corresponding experimental results.
Wenjun Li 0001, Zhenyu Chen 0003, Xingyu Gao 0001, Wei Liu 0010, Jin Wang 0001
IEEE Internet Things J.2
2017 BrainStorm: a psychosocial game suite design for non-invasive cross-generational cognitive capabilities data collection
abstract
Currently available traditional as well as videogame-based cognitive assessment techniques are inappropriate due to several reasons. This paper presents a novel psychosocial game suite, BrainStorm, for non-invasive cross-generational cognitive capabilities data collection, which additionally provides cross-generational social support. A motivation behind the development of presented game suite is to provide an entertaining and exciting platform for its target users in order to collect gameplay-based cognitive capabilities data in a non-invasive manner. An extensive evaluation of the presented game suite demonstrated high acceptability and attraction for its target users. Besides, the data collection process is successfully reported as transparent and non-invasive.
Yiqiang Chen 0001, Lisha Hu, Shuangquan Wang, Jindong Wang 0001, Zhenyu Chen 0003, Xinlong Jiang, Jianfei Shen
J. Exp. Theor. Artif. Intell.6
2017 Sparse Online Learning of Image Similarity
abstract
Learning image similarity plays a critical role in real-world multimedia information retrieval applications, especially in Content-Based Image Retrieval (CBIR) tasks, in which an accurate retrieval of visually similar objects largely relies on an effective image similarity function. Crafting a good similarity function is very challenging because visual contents of images are often represented as feature vectors in high-dimensional spaces, for example, via bag-of-words (BoW) representations, and traditional rigid similarity functions, for example, cosine similarity, are often suboptimal for CBIR tasks. In this article, we address this fundamental problem, that is, learning to optimize image similarity with sparse and high-dimensional representations from large-scale training data, and propose a novel scheme of Sparse Online Learning of Image Similarity (SOLIS). In contrast to many existing image-similarity learning algorithms that are designed to work with low-dimensional data, SOLIS is able to learn image similarity from large-scale image data in sparse and high-dimensional spaces. Our encouraging results showed that the proposed new technique achieves highly competitive accuracy as compared to the state-of-the-art approaches but enjoys significant advantages in computational efficiency, model sparsity, and retrieval scalability, making it more practical for real-world multimedia retrieval applications.
Xingyu Gao 0001, Steven C. H. Hoi, Yongdong Zhang 0001, Jianshe Zhou, Ji Wan, Zhenyu Chen 0003, Jintao Li 0001, Jianke Zhu
ACM Trans. Intell. Syst. Technol.6
2016 A Study of Players' Experiences During Brain Games Play
Yiqiang Chen 0001, Shuangquan Wang, Zhenyu Chen 0003, Jianfei Shen, Lisha Hu, Jindong Wang 0001
PRICAI4
2016 Adaptive weighted imbalance learning with application to abnormal activity recognition
Xingyu Gao 0001, Zhenyu Chen 0003, Sheng Tang, Yongdong Zhang 0001, Jintao Li 0001
Neurocomputing2
2016 Feature Adaptive Online Sequential Extreme Learning Machine for lifelong indoor localization
Xinlong Jiang, Junfa Liu, Yiqiang Chen 0001, Dingjun Liu, Yang Gu 0001, Zhenyu Chen 0003
Neural Comput. Appl.6
2015 Unobtrusive Sensing Incremental Social Contexts Using Fuzzy Class Incremental Learning
abstract
By utilizing captured characteristics of surrounding contexts through widely used Bluetooth sensor, user-centric social contexts can be effectively sensed and discovered by dynamic Bluetooth information. At present, state-of-the-art approaches for building classifiers can basically recognize limited classes trained in the learning phase; however, due to the complex diversity of social contextual behavior, the built classifier seldom deals with newly appeared contexts, which results in degrading the recognition performance greatly. To address this problem, we propose, an OSELM (online sequential extreme learning machine) based class incremental learning method for continuous and unobtrusive sensing new classes of social contexts from dynamic Bluetooth data alone. We integrate fuzzy clustering technique and OSELM to discover and recognize social contextual behaviors by real-world Bluetooth sensor data. Experimental results show that our method can automatically cope with incremental classes of social contexts that appear unpredictably in the real-world. Further, our proposed method have the effective recognition capability for both original known classes and newly appeared unknown classes, respectively.
Zhenyu Chen 0003, Yiqiang Chen 0001, Xingyu Gao 0001, Shuangquan Wang, Lisha Hu, Chenggang Yan 0001, Nicholas D. Lane, Chunyan Miao
ICDM1
2014 StudentLife: assessing mental health, academic performance and behavioral trends of college students using smartphones
abstract
Much of the stress and strain of student life remains hidden. The StudentLife continuous sensing app assesses the day-to-day and week-by-week impact of workload on stress, sleep, activity, mood, sociability, mental well-being and academic performance of a single class of 48 students across a 10 week term at Dartmouth College using Android phones. Results from the StudentLife study show a number of significant correlations between the automatic objective sensor data from smartphones and mental health and educational outcomes of the student body. We also identify a Dartmouth term lifecycle in the data that shows students start the term with high positive affect and conversation levels, low stress, and healthy sleep and daily activity patterns. As the term progresses and the workload increases, stress appreciably rises while positive affect, sleep, conversation and activity drops off. The StudentLife dataset is publicly available on the web.
Rui Wang 0016, Zhenyu Chen 0003, Tianxing Li 0001, Gabriella M. Harari, Stefanie Tignor, Dror Ben-Zeev, Andrew T. Campbell
UbiComp3
2014 Boosting cross-media retrieval via visual-auditory feature analysis and relevance feedback
abstract
Different types of multimedia data express high-level semantics from different aspects. How to learn comprehensive high-level semantics from different types of data and enable efficient cross-media retrieval becomes an emerging hot issue. There are abundant statistical and semantic correlations among heterogeneous low-level media content, which makes it challenging to query cross-media data effectively. In this paper, we propose a new cross-media retrieval method based on short-term and long-term relevance feedback. Our method mainly focuses on two typical types of media data, i.e. image and audio. First, we build multimodal representation via statistical canonical correlation between image and audio feature matrices, and define cross-media distance metric for similarity measure; then we propose optimization strategy based on relevance feedback, which fuses short-term learning results and long-term accumulated knowledge into the objective function. Experiments on image-audio dataset have demonstrated the superiority of our method over several existing algorithms.
Hong Zhang 0022, Junsong Yuan 0001, Xingyu Gao 0001, Zhenyu Chen 0003
ACM Multimedia4
2014 b-COELM: A fast, lightweight and accurate activity recognition model for mini-wearable devices
Lisha Hu, Yiqiang Chen 0001, Shuangquan Wang, Zhenyu Chen 0003
Pervasive Mob. Comput.4
2013 CarSafe app: alerting drowsy and distracted drivers using dual cameras on smartphones
abstract
We present CarSafe, a new driver safety app for Android phones that detects and alerts drivers to dangerous driving conditions and behavior. It uses computer vision and machine learning algorithms on the phone to monitor and detect whether the driver is tired or distracted using the front-facing camera while at the same time tracking road conditions using the rear-facing camera. Today's smartphones do not, however, have the capability to process video streams from both the front and rear cameras simultaneously. In response, CarSafe uses acontext-aware algorithm that switches between the two cameras while processing the data in real-time with the goal of minimizing missed events inside (e.g., drowsy driving) and outside of the car (e.g., tailgating). Camera switching means that CarSafe technically has a "blind spot" in the front or rear at any given time. To address this, CarSafe uses other embedded sensors on the phone (i.e., inertial sensors) to generate soft hints regarding potential blind spot dangers. We present the design and implementation of CarSafe and discuss its evaluation using results from a 12-driver field trial. Results from the CarSafe deployment are promising -- CarSafe can infer a common set of dangerous driving behaviors and road conditions with an overall precision and recall of 83% and 75%, respectively. CarSafe is the first dual-camera sensing app for smartphones and represents a new disruptive technology because it provides similar advanced safety features otherwise only found in expensive top-end cars.
Chuang-Wen You, Nicholas D. Lane, Rui Wang 0016, Zhenyu Chen 0003, Thomas J. Bao, Martha Montes-de-Oca, Yuting Cheng 0001, Mu Lin, Lorenzo Torresani, Andrew T. Campbell
MobiSys5
2013 CarSafe app: alerting drowsy and distracted drivers using dual cameras on smartphones
abstract
We present CarSafe, the first driver safety application that uses dual cameras on smartphones to detect and alert drivers to dangerous driving conditions. CarSafe fuses events detected from cameras and readings from embedded sensors on the phone -- such as the GPS, accelerometer and gyroscope -- to detect and alert the driver of dangerous driving behavior in and outside of the car. Results from a 12-driver field trial show CarSafe can infer five of the most commonly occurring dangerous driving conditions with an overall precision and recall of 83% and 75%, respectively.
Chuang-Wen You, Nicholas D. Lane, Rui Wang 0016, Zhenyu Chen 0003, Thomas J. Bao, Martha Montes-de-Oca, Yuting Cheng 0001, Mu Lin, Lorenzo Torresani, Andrew T. Campbell
MobiSys5
2012 Surrounding context and episode awareness using dynamic Bluetooth data
abstract
Bluetooth information can efficiently capture characteristics of user-centric surrounding contexts, such as formal meeting or chatting with friends, shopping with friends or alone, etc. In this paper, we extract novel features from Bluetooth traces and use these features for recognizing contextual behavior as well as inferring continuous episode transition. Evaluation results show that extracted novel features are very effective, which enable the model to achieve an average of 87% accuracy for specific context classification and the ability of episode inference from real-life Bluetooth traces.
Yiqiang Chen 0001, Zhenyu Chen 0003, Junfa Liu, Derek Hao Hu, Qiang Yang 0001
UbiComp2
2012 Extreme learning machine-based device displacement free activity recognition model
Yiqiang Chen 0001, Zhongtang Zhao, Shuangquan Wang, Zhenyu Chen 0003
Soft Comput.4
2010 Automatic Generation of Pencil Sketch for 2D Images
Xingyu Gao 0001, Jingye Zhou, Zhenyu Chen 0003, Yiqiang Chen 0001
ICASSP3