VLDB 2026 Research / reviewers in the wild / expert
Peijun Zhao
dblp:166/6002
· DBLP profile ↗
26ranked-venue papers
7as first author
10since 2021 · last 2026
0000-0003-2843-6941ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 10 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 since 2021Systems, architecture and hardware · 3 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 3Security and privacy · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Physical Self-Supervised Learning: IMU Sensing without Manual LabelsabstractDeep neural networks have become a promising approach for IMU-based sensing, but their scalability is fundamentally limited by costly labeled data and poor robustness to heterogeneous devices, placements, and users. Existing unsupervised and self-supervised methods reduce but do not remove this dependence, still requiring labeled data for domain adaptation and largely ignoring known physical structure. We propose physical self-supervised learning, an autoencoder-style paradigm for label-free IMU sensing. We replace the conventional neural decoder with an auto-adaptive physics decoder—a learnable family of kinematic equations that enforces explicit physical structure while adapting across environments—and adopt a hybrid two-stage IMU encoder with reconstruction in a structured latent space to mitigate sensor noise. Our framework further introduces probabilistic frequency-spatial constraints to disentangle sensor and object motion, a multi-view kinematic tree to exploit sparse physical self-supervised signals, and an uncertainty-aware formulation to handle the inherent ambiguity of IMU inference. Evaluated on inertial tracking and full-body motion capture over public datasets and realistic deployments, physical self-supervised learning reduces errors by up to 5× for tracking and 4× for motion capture in challenging generalization scenarios, consistently outperforming state-of-the-art supervised and self-supervised baselines without any labels. Yuyang Leng, Renyuan Liu, Shaohan Hu, Peijun Zhao, Chun-Fu Chen 0001, Songqing Chen, Shuochao Yao |
MobiSys | 4 |
| 2025 | MSFS: Maliciously Secure 3-Party Feature Selection via Mutual Information
Peijun Zhao, Lin Liu 0018, Shaojing Fu, Yuchuan Luo |
Inscrypt (2) | 1 |
| 2025 | DAF: An Efficient End-to-End Dynamic Activation Framework for on-Device DNN TrainingabstractRecent advancements in on-device training for deep neural networks have underscored the critical need for efficient activation compression to overcome the memory constraints of mobile and edge devices. As activations dominate memory usage during training and are essential for gradient computation, compressing them without compromising accuracy remains a key research challenge. While existing methods for dynamic activation quantization promise theoretical memory savings, their practical deployment is impeded by system-level challenges such as computational overhead and memory fragmentation. Renyuan Liu, Yuyang Leng, Kaiyan Liu, Shaohan Hu, Chun-Fu Chen 0001, Peijun Zhao, Heechul Yun, Shuochao Yao |
MobiSys | 6 |
| 2024 | milliFlow: Scene Flow Estimation on mmWave Radar Point Cloud for Human Motion Sensing
Fangqiang Ding, Peijun Zhao, Xiaoxuan Lu 0001 |
ECCV (24) | 3 |
| 2023 | mmPoint: Dense Human Point Cloud Generation from mmWave
Qian Xie 0001, Qianyi Deng, Ta Ying Cheng, Peijun Zhao, Amir Patel, Agathoniki Trigoni, Andrew Markham |
BMVC | 4 |
| 2023 | Motor Function Assessment of Children with Cerebral Palsy using Monocular VideoabstractThe assessment of movement abilities in individuals with neurological disorders is a critical task in clinical practice. Currently, clinical assessments are time-consuming and rely on qualitative scales typically conducted by trained clinicians. Moreover, these assessments offer only coarse snapshots of a person’s abilities, failing to track the minutiae of recovery over time. To overcome these limitations, we propose a machine learning approach based on spatial-temporal graph convolutional network (STGCN) to extract movement features from pose data obtained from monocular videos collected with mobile devices (e.g., smartphones, tablets). Our proposed method achieves an accuracy of 76% in evaluating children with Cerebral Palsy (CP) using the Gross Motor Function Classification System (GMFCS), a 10% improvement in accuracy compared to current state-of-the-art methods, and shows substantial agreement with professional assessments based on the weighted Cohen’s Kappa (κlw= 0.74). Furthermore, the proposed method can be efficiently implemented on a wide range of mobile devices in real-time or near real-time. Peijun Zhao, Moises Alencastre-Miranda, Zhan Shen, Ciaran O'Neill, David Whiteman, Javier Gervas-Arruga, Hermano Igo Krebs |
BSN | 1 |
| 2023 | CubeLearn: End-to-End Learning for Human Motion Recognition From Raw mmWave Radar SignalsabstractmmWave FMCW radar has attracted a huge amount of research interest for human-centered applications in recent years, such as human gesture and activity recognition. Most existing pipelines are built upon conventional discrete Fourier transform (DFT) preprocessing and deep neural network classifier hybrid methods, with a majority of previous works focusing on designing the downstream classifier to improve overall accuracy. In this work, we take a step back and look at the preprocessing module. To avoid the drawbacks of conventional DFT preprocessing, we propose a complex-weighted learnable preprocessing module, named CubeLearn, to directly extract features from raw radar signal and build an end-to-end deep neural network for mmWave FMCW radar motion recognition applications. Extensive experiments show that our CubeLearn module consistently improves the classification accuracies of different pipelines, especially, benefiting those simpler models, which are more likely to be used on edge devices due to their computational efficiency. We provide ablation studies on initialization methods and structure of the proposed module, as well as an evaluation of the running time on PC and edge devices. This work also serves as a comparison of different approaches toward data cube slicing. Through our task-agnostic design, we propose a first step toward a generic end-to-end solution for radar recognition problems. Peijun Zhao, Xiaoxuan Lu 0001, Bing Wang 0013, Agathoniki Trigoni, Andrew Markham |
IEEE Internet Things J. | 1 |
| 2021 | P2-Net: Joint Description and Detection of Local Features for Pixel and Point MatchingabstractAccurately describing and detecting 2D and 3D key-points is crucial to establishing correspondences across images and point clouds. Despite a plethora of learning-based 2D or 3D local feature descriptors and detectors having been proposed, the derivation of a shared descriptor and joint keypoint detector that directly matches pixels and points remains under-explored by the community. This work takes the initiative to establish fine-grained correspondences between 2D images and 3D point clouds. In order to directly match pixels and points, a dual fully-convolutional framework is presented that maps 2D and 3D inputs into a shared latent representation space to simultaneously describe and detect keypoints. Furthermore, an ultra-wide reception mechanism and a novel loss function are designed to mitigate the intrinsic information variations between pixel and point local regions. Extensive experimental results demonstrate that our framework shows competitive performance in fine-grained matching between images and point clouds and achieves state-of-the-art results for the task of indoor visual localization. Our source code is available at https://github.com/BingCS/P2-Net. Bing Wang 0013, Changhao Chen, Zhaopeng Cui, Jie Qin 0004, Xiaoxuan Lu 0001, Zhengdi Yu, Peijun Zhao, Zhen Dong 0005, Fan Zhu 0001, Agathoniki Trigoni, Andrew Markham |
ICCV | 7 |
| 2021 | 3D Motion Capture of an Unmodified Drone with Single-chip Millimeter Wave RadarabstractAccurate motion capture of aerial robots in 3D is a key enabler for autonomous operation in indoor environments such as warehouses or factories, as well as driving forward research in these areas. The most commonly used solutions at present are optical motion capture (e.g. VICON) and Ultrawide-band (UWB), but these are costly and cumbersome to deploy, due to their requirement of multiple cameras/anchors spaced around the tracking area. They also require the drone to be modified to carry an active or passive marker. In this work, we present an inexpensive system that can be rapidly installed, based on single-chip millimeter wave (mmWave) radar. Importantly, the drone does not need to be modified or equipped with any markers, as we exploit the Doppler signals from the rotating propellers. Furthermore, 3D tracking is possible from a single point, greatly simplifying deployment. We develop a novel deep neural network and demonstrate decimeter level 3D tracking at 10Hz, achieving better performance than classical baselines. Our hope is that this low-cost system will act to catalyse inexpensive drone research and increased autonomy. Peijun Zhao, Xiaoxuan Lu 0001, Bing Wang 0013, Agathoniki Trigoni, Andrew Markham |
ICRA | 1 |
| 2021 | Human tracking and identification through a millimeter wave radar
Peijun Zhao, Xiaoxuan Lu 0001, Changhao Chen, Wei Wang 0226, Agathoniki Trigoni, Andrew Markham |
Ad Hoc Networks | 1 |
| 2020 | AtLoc: Attention Guided Camera LocalizationabstractDeep learning has achieved impressive results in camera localization, but current single-image techniques typically suffer from a lack of robustness, leading to large outliers. To some extent, this has been tackled by sequential (multi-images) or geometry constraint approaches, which can learn to reject dynamic objects and illumination conditions to achieve better performance. In this work, we show that attention can be used to force the network to focus on more geometrically robust objects and features, achieving state-of-the-art performance in common benchmark, even if using only a single image as input. Extensive experimental evidence is provided through public indoor and outdoor datasets. Through visualization of the saliency maps, we demonstrate how the network learns to reject dynamic objects, yielding superior global camera pose regression performance. The source code is avaliable at https://github.com/BingCS/AtLoc. Bing Wang 0013, Changhao Chen, Xiaoxuan Lu 0001, Peijun Zhao, Agathoniki Trigoni, Andrew Markham |
AAAI | 4 |
| 2020 | Heart Rate Sensing with a Robot Mounted mmWave RadarabstractHeart rate monitoring at home is a useful metric for assessing health e.g. of the elderly or patients in post-operative recovery. Although non-contact heart rate monitoring has been widely explored, typically using a static, wall-mounted device, measurements are limited to a single room and sensitive to user orientation and position. In this work, we propose mBeats, a robot mounted millimeter wave (mmWave) radar system that provide periodic heart rate measurements under different user poses, without interfering in a users daily activities. mBeats contains a mmWave servoing module that adaptively adjusts the sensor angle to the best reflection pro le. Furthermore, mBeats features a deep neural network predictor, which can estimate heart rate from the lower leg and additionally provides estimation uncertainty. Through extensive experiments, we demonstrate accurate and robust operation of mBeats in a range of scenarios. We believe by integrating mobility and adaptability, mBeats can empower many down-stream healthcare applications at home, such as palliative care, post-operative rehabilitation and telemedicine. Peijun Zhao, Xiaoxuan Lu 0001, Bing Wang 0013, Changhao Chen, Linhai Xie, Agathoniki Trigoni, Andrew Markham |
ICRA | 1 |
| 2020 | See through smoke: robust indoor mapping with low-cost mmWave radarabstractThis paper presents the design, implementation and evaluation of milliMap, a single-chip millimetre wave (mmWave) radar based indoor mapping system targetted towards low-visibility environments to assist in emergency response. A unique feature of milliMap is that it only leverages a low-cost, off-the-shelf mmWave radar, but can reconstruct a dense grid map with accuracy comparable to lidar, as well as providing semantic annotations of objects on the map. milliMap makes two key technical contributions. First, it autonomously overcomes the sparsity and multi-path noise of mmWave signals by combining cross-modal supervision from a co-located lidar during training and the strong geometric priors of indoor spaces. Second, it takes the spectral response of mmWave reflections as features to robustly identify different types of objects e.g. doors, walls etc. Extensive experiments in different indoor environments show that milliMap can achieve a map reconstruction error less than 0.2m and classify key semantics with an accuracy of ~ 90%, whilst operating through dense smoke. Xiaoxuan Lu 0001, Stefano Rosa, Peijun Zhao, Bing Wang 0013, Changhao Chen, John A. Stankovic, Agathoniki Trigoni, Andrew Markham |
MobiSys | 3 |
| 2020 | milliEgo: single-chip mmWave radar aided egomotion estimation via deep sensor fusionabstractRobust and accurate trajectory estimation of mobile agents such as people and robots is a key requirement for providing spatial awareness for emerging capabilities such as augmented reality or autonomous interaction. Although currently dominated by optical techniques e.g., visual-inertial odometry these suffer from challenges with scene illumination or featureless surfaces. As an alternative, we propose milliEgo, a novel deep-learning approach to robust egomotion estimation which exploits the capabilities of low-cost mm Wave radar. Although mmWave radar has a fundamental advantage over monocular cameras of being metric i.e., providing absolute scale or depth, current single chip solutions have limited and sparse imaging resolution, making existing point-cloud registration techniques brittle. We propose a new architecture that is optimized for solving this challenging pose transformation problem. Secondly, to robustly fuse mmWave pose estimates with additional sensors, e.g. inertial or visual sensors we introduce a mixed attention approach to deep fusion. Through extensive experiments, we demonstrate our proposed system is able to achieve 1.3% 3D error drift and generalizes well to unseen environments. We also show that the neural architecture can be made highly efficient and suitable for real-time embedded applications. Xiaoxuan Lu 0001, Muhamad Risqi Utama Saputra, Peijun Zhao, Yasin Almalioglu, Pedro Porto Buarque de Gusmão, Changhao Chen, Ke Sun 0012, Agathoniki Trigoni, Andrew Markham |
SenSys | 3 |
| 2020 | Deep-Learning-Based Pedestrian Inertial Navigation: Methods, Data Set, and On-Device InferenceabstractModern inertial measurements units (IMUs) are small, cheap, energy efficient, and widely employed in smart devices and mobile robots. Exploiting inertial data for accurate and reliable pedestrian navigation supports is a key component for emerging Internet of Things applications and services. Recently, there has been a growing interest in applying deep neural networks (DNNs) to motion sensing and location estimation. However, the lack of sufficient labelled data for training and evaluating architecture benchmarks has limited the adoption of DNNs in IMU-based tasks. In this article, we present and release the Oxford Inertial Odometry Data Set (OxIOD), a first-of-its-kind public data set for deep-learning-based inertial navigation research with fine-grained ground truth on all sequences. Furthermore, to enable more efficient inference at the edge, we propose a novel lightweight framework to learn and reconstruct pedestrian trajectories from raw IMU data. Extensive experiments show the effectiveness of our data set and methods in achieving accurate data-driven pedestrian inertial navigation on resource-constrained devices. Changhao Chen, Peijun Zhao, Xiaoxuan Lu 0001, Wei Wang 0226, Andrew Markham, Agathoniki Trigoni |
IEEE Internet Things J. | 2 |
| 2020 | High-quality face image generation based on generative adversarial networksabstractConventional face image generation using generative adversarial networks (GAN) is limited by the quality of generated images since generator and discriminator use the same backpropagation network . In this paper, we discuss algorithms that can improve the quality of generated images, that is, high-quality face image generation . In order to achieve stability of network, we replace MLP with convolutional neural network (CNN) and remove pooling layers. We conduct comprehensive experiments on LFW, CelebA datasets and experimental results show the effectiveness of our proposed method. Zhixin Zhang 0005, Xuhua Pan, Shuhao Jiang, Peijun Zhao |
J. Vis. Commun. Image Represent. | 4 |
| 2019 | mID: Tracking and Identifying People with Millimeter Wave RadarabstractThe key to offering personalised services in smart spaces is knowing where a particular person is with a high degree of accuracy. Visual tracking is one such solution, but concerns arise around the potential leakage of raw video information and many people are not comfortable accepting cameras in their homes or workplaces. We propose a human tracking and identification system (mID) based on millimeter wave radar which has a high tracking accuracy, without being visually compromising. Unlike competing techniques based on WiFi Channel State Information (CSI), it is capable of tracking and identifying multiple people simultaneously. Using a lowcost, commercial, off-the-shelf radar, we first obtain sparse point clouds and form temporally associated trajectories. With the aid of a deep recurrent network, we identify individual users. We evaluate and demonstrate our system across a variety of scenarios, showing median position errors of 0.16 m and identification accuracy of 89% for 12 people. Peijun Zhao, Xiaoxuan Lu 0001, Changhao Chen, Wei Wang 0226, Agathoniki Trigoni, Andrew Markham |
DCOSS | 1 |
| 2019 | DeepPCO: End-to-End Point Cloud Odometry through Deep Parallel Neural NetworkabstractOdometry is of key importance for localization in the absence of a map. There is considerable work in the area of visual odometry (VO), and recent advances in deep learning have brought novel approaches to VO, which directly learn salient features from raw images. These learning-based approaches have led to more accurate and robust VO systems. However, they have not been well applied to point cloud data yet. In this work, we investigate how to exploit deep learning to estimate point cloud odometry (PCO), which may serve as a critical component in point cloud-based downstream tasks or learning-based systems. Specifically, we propose a novel end-to-end deep parallel neural network called DeepPCO, which can estimate the 6-DOF poses using consecutive point clouds. It consists of two parallel sub-networks to estimate 3D translation and orientation respectively rather than a single neural network. We validate our approach on KITTI Visual Odometry/SLAM benchmark dataset with different baselines. Experiments demonstrate that the proposed approach achieves good performance in terms of pose accuracy. Wei Wang 0226, Muhamad Risqi Utama Saputra, Peijun Zhao, Pedro Porto Buarque de Gusmão, Bo Yang 0027, Changhao Chen, Andrew Markham, Agathoniki Trigoni |
IROS | 3 |
| 2019 | Autonomous Learning of Speaker Identity and WiFi Geofence From Noisy Sensor DataabstractA fundamental building block toward intelligent environments is the ability to understand who is present in a certain area. A ubiquitous way of detecting this is to exploit unique vocal characteristics as people interact with one another in common spaces. However, manually enrolling users into a biometric database is time-consuming and not robust to vocal deviations over time. Instead, consider audio features sampled during a meeting, yielding a noisy set of possible voiceprints. With a number of meetings and knowledge of participation, e.g., sniffed wireless media access control (MAC) addresses, can we learn to associate a specific identity with a particular voiceprint? To address this problem, this paper advocates an Internet of Things (IoT) solution and proposes to use co-located WiFi as supervisory weak labels to automatically bootstrap the labeling process. In particular, a novel cross-modality labeling algorithm is proposed that jointly optimizes the clustering and association process, which solves the inherent mismatching issues arising from heterogeneous sensor data. At the same time, we further propose to reuse the labeled data to iteratively update wireless geofence models and curate device specific thresholds. The extensive experimental results from two different scenarios demonstrate that our proposed method is able to achieve twofold improvement in labeling compared with conventional methods and can achieve reliable speaker recognition in the wild. Xiaoxuan Lu 0001, Yuanbo Xiangli, Peijun Zhao, Changhao Chen, Agathoniki Trigoni, Andrew Markham |
IEEE Internet Things J. | 3 |
| 2018 | Deepauth: in-situ authentication for smartwatches via deeply learned behavioural biometricsabstractThis paper proposes DeepAuth, an in-situ authentication framework that leverages the unique motion patterns when users entering passwords as behavioural biometrics. It uses a deep recurrent neural network to capture the subtle motion signatures during password input, and employs a novel loss function to learn deep feature representations that are robust to noise, unseen passwords, and malicious imposters even with limited training data. DeepAuth is by design optimised for resource constrained platforms, and uses a novel split-RNN architecture to slim inference down to run in real-time on off-the-shelf smartwatches. Extensive experiments with real-world data show that DeepAuth outperforms the state-of-the-art significantly in both authentication performance and cost, offering real-time authentication on a variety of smartwatches. Xiaoxuan Lu 0001, Bowen Du 0002, Peijun Zhao, Hongkai Wen 0001, Yiran Shen 0001, Andrew Markham, Agathoniki Trigoni |
UbiComp | 3 |
| 2018 | Simultaneous Localization and Mapping with Power Network Electromagnetic FieldabstractVarious sensing modalities have been exploited for indoor location sensing, each of which has well understood limitations, however. This paper presents a first systematic study on using the electromagnetic field (EMF) induced by a building's electric power network for simultaneous localization and mapping (SLAM). A basis of this work is a measurement study showing that the power network EMF sensed by either a customized sensor or smartphone's microphone as a side-channel sensor is spatially distinct and temporally stable. Based on this, we design a SLAM approach that can reliably detect loop closures based on EMF sensing results. With the EMF feature map constructed by SLAM, we also design an efficient online localization scheme for resource-constrained mobiles. Evaluation in three indoor spaces shows that the power network EMF is a promising modality for location sensing on mobile devices, which is able to run in real time and achieve sub-meter accuracy. Xiaoxuan Lu 0001, Yang Li 0147, Peijun Zhao, Changhao Chen, Linhai Xie, Hongkai Wen 0001, Rui Tan 0001, Agathoniki Trigoni |
MobiCom | 3 |
| 2018 | Automatic Face Recognition Adaptation via Ambient Wireless IdentifiersabstractFace recognition is a key enabling service for smart-spaces, allowing building management agents to easily monitor 'who is where', anticipating user needs and tailoring their local environment and experiences. Although facial recognition, especially through the use of deep neural networks, has achieved stellar performance over large datasets, the majority of approaches require supervised learning, that is, to be trained with tens or hundreds of images of users in different poses and lighting conditions. In this paper, we motivate that this enrollment effort is unnecessary if the smart-space has access to a wireless identifier e.g., through a smart-phone's MAC address. By learning and refining the noisy and weak association between a user's smart-phone and facial images, AutoTune can fine-tune a deep neural network to tailor it to the environment, users and conditions of a particular camera or set of cameras. Xiaoxuan Lu 0001, Peijun Zhao, Bowen Du 0002, Hongkai Wen 0001, Andrew Markham, Stefano Rosa, Agathoniki Trigoni |
SenSys | 2 |
| 2017 | ViVo: Video-Augmented Dictionary for Vocabulary LearningabstractResearch on Computer-Assisted Language Learning (CALL) has shown that the use of multimedia materials such as images and videos can facilitate interpretation and memorization of new words and phrases by providing richer cues than text alone. We present ViVo, a novel video-augmented dictionary that provides an inexpensive, convenient, and scalable way to exploit huge online video resources for vocabulary learning. ViVo automatically generates short video clips from existing movies with the target word highlighted in the subtitles. In particular, we apply a word sense disambiguation algorithm to identify the appropriate movie scenes with adequate contextual information for learning. We analyze the challenges and feasibility of this approach and describe our interaction design. A user study showed that learners were able to retain nearly 30% more new words with ViVo than with a standard bilingual dictionary days after learning. They preferred our video-augmented dictionary for its benefits in memorization and enjoyable learning experience. Yeshuang Zhu, Yuntao Wang 0001, Chun Yu, Shaoyun Shi, Yankai Zhang, Shuang He, Peijun Zhao, Xiaojuan Ma, Yuanchun Shi |
CHI | 7 |
| 2017 | Inferring Emotional Tags From Social Images With User DemographicsabstractSocial images, which are images uploaded and shared on social networks, are used to express users’ emotions. Inferring emotional tags from social images is of great importance; it can benefit many applications, such as image retrieval and recommendation. Whereas previous related research has primarily focused on exploring image visual features, we aim to address this problem by studying whether user demographics make a difference regarding users’ emotional tags of social images. We first consider how to model the emotions of social images. Then, we investigate how user demographics, such as gender, marital status, and occupation, are related to the emotional tags of social images. A partially labeled factor graph model named the demographics factor graph model ( D-FGM ) is proposed to leverage the uncovered patterns. Experiments on a data set collected from the world's largest image sharing website Flickr 1 1 [Online]. Available: http://www.flickr.com/ confirm the accuracy of the proposed model. We also find some interesting phenomena. For example, men and women have different patterns to tag “anger” for social images. Boya Wu, Jia Jia 0001, Yang Yang 0009, Peijun Zhao, Jie Tang 0001, Qi Tian 0001 |
IEEE Trans. Multim. | 4 |
| 2016 | One-Dimensional Handwriting: Inputting Letters and Words on Smart GlassesabstractWe present 1D Handwriting, a unistroke gesture technique enabling text entry on a one-dimensional interface. The challenge is to map two-dimensional handwriting to a reduced one-dimensional space, while achieving a balance between memorability and performance efficiency. After an iterative design, we finally derive a set of ambiguous two-length unistroke gestures, each mapping to 1-4 letters. To input words, we design a Bayesian algorithm that takes into account the probability of gestures and the language model. To input letters, we design a pause gesture allowing users to switch into letter selection mode seamlessly. Users studies show that 1D Handwriting significantly outperforms a selection-based technique (a variation of 1Line Keyboard) for both letter input (4.67 WPM vs. 4.20 WPM) and word input (9.72 WPM vs. 8.10 WPM). With extensive training, text entry rate can reach 19.6 WPM. Users' subjective feedback indicates 1D Handwriting is easy to learn and efficient to use. Moreover, it has several potential applications for other one-dimensional constrained interfaces. Chun Yu, Ke Sun 0003, Mingyuan Zhong 0001, Peijun Zhao, Yuanchun Shi |
CHI | 5 |
| 2015 | Understanding the emotions behind social images: Inferring with user demographicsabstractUnderstanding the essential emotions behind social images is of vital importance: it can benefit many applications such as image retrieval and personalized recommendation. While previous related research mostly focuses on the image visual features, in this paper, we aim to tackle this problem by “linking inferring with users' demographics”. Specifically, we propose a partially-labeled factor graph model named D-FGM, to predict the emotions embedded in social images not only by the image visual features, but also by the information of users' demographics. We investigate whether users' demographics like gender, marital status and occupation are related to emotions of social images, and then leverage the uncovered patterns into modeling as different factors. Experiments on a data set from the world's largest image sharing website Flickr1 confirm the accuracy of the proposed model. The effectiveness of the users' demographics factors is also verified by the factor contribution analysis, which reveals some interesting behavioral phenomena as well. Boya Wu, Jia Jia 0001, Yang Yang 0009, Peijun Zhao, Jie Tang 0001 |
ICME | 4 |