Jingyi Hou

dblp:173/4937 · DBLP profile ↗
← Back
12ranked-venue papers
7as first author
5since 2021 · last 2025
0000-0003-0172-3115ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 6 first-authorArtificial intelligence and machine learning · 5 · 3 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Video understanding and tracking · 36% Knowledge representation and reasoning · 18% Vision and language · 16%

Topics — the 15 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Video understanding and tracking
action recognition
1.132020
Confidence-Guided Self Refinement for Action Prediction in Untrimmed Videos · IEEE Trans. Image Process. 2020
Content-Attention Representation by Factorized Action-Scene Network for Action Recognition · IEEE Trans. Multim. 2018
Unsupervised Deep Learning of Mid-Level Video Representation for Action Recognition · AAAI 2018
Computer vision › Video understanding and tracking
action anticipation
0.922021
Spatial-Temporal Relation Reasoning for Action Prediction in Videos · Int. J. Comput. Vis. 2021
Confidence-Guided Self Refinement for Action Prediction in Untrimmed Videos · IEEE Trans. Image Process. 2020
Computer vision › Vision and language
video captioning
0.822020
Joint Commonsense and Relation Reasoning for Image and Video Captioning · AAAI 2020
Joint Syntax Representation Learning and Visual Cue Translation for Video Captioning · ICCV 2019
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
latent factor model
0.812024
Discovering Predictable Latent Factors for Time Series Forecasting · IEEE Trans. Knowl. Data Eng. 2024
Machine learning › Time series and sequential data › time series analysis
time series forecasting
0.812024
Discovering Predictable Latent Factors for Time Series Forecasting · IEEE Trans. Knowl. Data Eng. 2024
Knowledge, reasoning and agents › Knowledge representation and reasoning › temporal reasoning
spatio-temporal reasoning
0.512021
Spatial-Temporal Relation Reasoning for Action Prediction in Videos · Int. J. Comput. Vis. 2021
Knowledge, reasoning and agents › Knowledge representation and reasoning
commonsense reasoning
0.412020
Joint Commonsense and Relation Reasoning for Image and Video Captioning · AAAI 2020
Computer vision › Vision and language
image captioning
0.412020
Joint Commonsense and Relation Reasoning for Image and Video Captioning · AAAI 2020
Knowledge, reasoning and agents › Knowledge representation and reasoning
relational reasoning
0.412020
Joint Commonsense and Relation Reasoning for Image and Video Captioning · AAAI 2020
Computer vision › Video understanding and tracking › action recognition
spatio-temporal action recognition
0.412020
Confidence-Guided Self Refinement for Action Prediction in Untrimmed Videos · IEEE Trans. Image Process. 2020
Computer vision › Video understanding and tracking › event recognition
complex event detection
0.312018
Content-Attention Representation by Factorized Action-Scene Network for Action Recognition · IEEE Trans. Multim. 2018
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
discriminative clustering
0.312018
Unsupervised Deep Learning of Mid-Level Video Representation for Action Recognition · AAAI 2018
Machine learning › Learning paradigms › unsupervised learning
unsupervised video representation learning
0.312018
Unsupervised Deep Learning of Mid-Level Video Representation for Action Recognition · AAAI 2018
Machine learning › Deep learning architectures and training
attention mechanism
0.112018
Content-Attention Representation by Factorized Action-Scene Network for Action Recognition · IEEE Trans. Multim. 2018
Machine learning › Deep learning architectures and training › attention mechanism
content attention
0.112018
Content-Attention Representation by Factorized Action-Scene Network for Action Recognition · IEEE Trans. Multim. 2018

Methods — techniques the papers use, named apart from their topics

transformer · 0.8deep latent dynamics models · 0.8relation reasoning · 0.5sparse self-attention · 0.4semantic graph · 0.4self-refining network · 0.4iterative learning · 0.4gumbel-softmax · 0.4confidence learning · 0.4commonsense reasoning · 0.4
YearPublicationVenuePosition
2025 Image sentiment analysis based on distillation and sentiment region localization network
abstract
Abstract Accurately identifying the emotions in images is crucial for sentiment content analysis. To detect local sentiment regions and acquire discriminative sentiment features, we propose a novel model named Distillation-guided and Contrastive-enhanced Sentiment Region Localization Network (DC-SRLN) to effectively complete image sentiment analysis. Two smart but heterogeneous SRLNs are designed first to pursue local sentiment regions. Then an innovative contrastive learning mode is implemented between global and local features to further enhance the discriminative ability of the sentiment features. Third, the enhanced global and local sentiment features are seamlessly integrated to guide each SRLN accurately capture local sentiment regions. Finally, an adaptive feature fusion module is created to fuse the heterogeneous features from the two SRLNs and generate a new multi-view multi-granularity sentiment semantics with more discriminative ability for image sentiment analysis. Extensive experimental results on three prevailing datasets, namely Twitter I, FI, and ArtPhoto, exhibit that DC-SRLN achieves satisfactory accuracies of 93.2%, 80.6%, and 78.7%, respectively, outperforming recent state-of-the-art baselines. Moreover, DC-SRLN needs less training time, demonstrating its high practicality. The code of DC-SRLN is freely available at https://github.com/Riley6868/DC-SRLN.
Hongbin Zhang 0004, Ya Feng, Jingyi Hou, Guangli Li
Comput. J.4
2024 Type-adaptive graph Transformer for heterogeneous information networks
Yanzhe Huang, Jingyi Hou, Zhijie Liu 0001
Appl. Intell.3
2024 Discovering Predictable Latent Factors for Time Series Forecasting
abstract
Modern temporal modeling methods, such as Transformer and its variants, have demonstrated remarkable capabilities in handling sequential data from specific domains like language and vision. Though achieving high performance with large-scale data, they often have redundant or unexplainable structures. When encountering some real-world datasets with limited observable variables that can be affected by many unknown factors, these methods may struggle to identify meaningful patterns and dependencies inherent in data, and thus, the modeling becomes unstable and unpredictable. To tackle this critical issue, in this article, we develop a novel algorithmic framework for inferring latent factors implied by the observed temporal data. The inferred factors are used to form multiple predictable and independent signal components that enable not only the reconstruction of future time series for accurate prediction but also sparse relation reasoning for long-term efficiency. To achieve this, we introduce three characteristics, i.e., predictability, sufficiency, and identifiability, and model these characteristics of latent factors via powerful deep latent dynamics models to infer the predictable signal components. Empirical results on multiple real datasets show the efficiency of our method for different kinds of time series forecasting tasks. Statistical analyses validate the predictability and interpretability of the learned latent factors.
Jingyi Hou, Zhen Dong 0002, Zhijie Liu 0001
IEEE Trans. Knowl. Data Eng.1
2023 Edge AI as a Service: Configurable Model Deployment and Delay-Energy Optimization With Result Quality Constraints
abstract
The breakthrough of artificial intelligence (AI) techniques has accelerated their applications in a wide range of industries, such as security protection, transportation, agriculture, and medical care. With the support of edge computing environments, providing latency guaranteed AI as a Service (AIaaS) can accelerate the deployment of data-intensive and computation-intensive AI applications and reduce the investment cost of the customers. However, the deployment architecture and working mechanism design, and performance optimization problems specific for AIaaS with configurable data quality and model complexity have not been studied in existing works. To address the problem, we propose a configurable model deployment architecture (CMDA) for edge AIaaS and present a flexible working mechanism by enabling the joint configuration of data quality ratios (DQRs) and model complexity ratios (MCRs) for the AI tasks. Along with commonly used resource allocation operations, the manager can improve the energy and delay performance of AI services with the desired quality of results (QoRs). We develop an energy-delay minimization problem under the framework of CMDA and propose a polynomial regression based relaxing method to solve the task configuration subproblem. We conduct experiments and simulations on the ImageNet classification and the common objects in context (COCO) object detection tasks using state-of-the-art deep learning models. We present the corresponding result quality tables (RQTs) and QoR regression models to illustrate the proposed method. The results of single task configuration and multi-task configuration and resource allocation on ImageNet classification and COCO object detection tasks demonstrate that the proposed method can achieve over$5\times$HDEC improvement compared with non-optimization schemes, and also show that joint configuration of DQR and MCR can achieve over$1.2\times$HDEC improvement compared with the methods that only configure DQR or MCR.
Wenyu Zhang 0002, Sherali Zeadally, Wei Li 0074, Haijun Zhang 0001, Jingyi Hou, Victor C. M. Leung
IEEE Trans. Cloud Comput.5
2021 Spatial-Temporal Relation Reasoning for Action Prediction in Videos
Xinxiao Wu, Jingyi Hou, Hanxi Lin, Jiebo Luo 0001
Int. J. Comput. Vis.3
2020 Joint Commonsense and Relation Reasoning for Image and Video Captioning
abstract
Exploiting relationships between objects for image and video captioning has received increasing attention. Most existing methods depend heavily on pre-trained detectors of objects and their relationships, and thus may not work well when facing detection challenges such as heavy occlusion, tiny-size objects, and long-tail classes. In this paper, we propose a joint commonsense and relation reasoning method that exploits prior knowledge for image and video captioning without relying on any detectors. The prior knowledge provides semantic correlations and constraints between objects, serving as guidance to build semantic graphs that summarize object relationships, some of which cannot be directly perceived from images or videos. Particularly, our method is implemented by an iterative learning algorithm that alternates between 1) commonsense reasoning for embedding visual regions into the semantic space to build a semantic graph and 2) relation reasoning for encoding semantic graphs to generate sentences. Experiments on several benchmark datasets validate the effectiveness of our prior knowledge-based approach.
Jingyi Hou, Xinxiao Wu, Xiaoxun Zhang, Yayun Qi, Yunde Jia, Jiebo Luo 0001
AAAI1
2020 Confidence-Guided Self Refinement for Action Prediction in Untrimmed Videos
abstract
Many existing methods formulate the action prediction task as recognizing early parts of actions in trimmed videos. In this paper, we focus on predicting actions from ongoing untrimmed videos where actions might not happen at the very beginning of videos. It is extremely challenging to predict actions in such untrimmed videos due to ambiguous or even no information of actions in the early parts of videos. To address this problem, we propose a prediction confidence that assesses the decision quality of a prediction model. Guided by the confidence, the model continuously refines the prediction results by itself with the increasing observed video frames. Specifically, we build a Self Prediction Refining Network (SPR-Net) which incrementally learns the confidence for action prediction. SPR-Net consists of three modules: a temporal hybrid network, an incremental confidence learner, and a self-refining Gumbel softmax sampler. The temporal hybrid network generates the action category distributions by integrating static scene and dynamic motion information. The incremental confidence learner calculates the confidence in an incremental manner, judging the extent to which the temporal hybrid network should believe its prediction result. The self-refining Gumbel softmax sampler models the mutual relationship between the prediction confidence and the category distribution, which enables them to be jointly learned in an end-to-end fashion. We also present a sparse self-attention mechanism to encode local spatio-temporal features into the frame-level motion representation to further improve the prediction performance. Extensive experiments on five datasets (i.e., UT-Interaction, BIT-Interaction, UCF101, THUMOS14, and ActivityNet) validate the effectiveness of the proposed method.
Jingyi Hou, Xinxiao Wu, Jiebo Luo 0001, Yunde Jia
IEEE Trans. Image Process.1
2019 Joint Syntax Representation Learning and Visual Cue Translation for Video Captioning
abstract
Video captioning is a challenging task that involves not only visual perception but also syntax representation learning. Recent progress in video captioning has been achieved through visual perception, but syntax representation learning is still under-explored. We propose a novel video captioning approach that takes into account both visual perception and syntax representation learning to generate accurate descriptions of videos. Specifically, we use sentence templates composed of Part-of-Speech (POS) tags to represent the syntax structure of captions, and accordingly, syntax representation learning is performed by directly inferring POS tags from videos. The visual perception is implemented by a mixture model which translates visual cues into lexical words that are conditional on the learned syntactic structure of sentences. Thus, a video captioning task consists of two sub-tasks: video POS tagging and visual cue translation, which are jointly modeled and trained in an end-to-end fashion. Evaluations on three public benchmark datasets demonstrate that our proposed method achieves substantially better performance than the state-of-the-art methods, which validates the superiority of joint modeling of syntax representation learning and visual perception for video captioning.
Jingyi Hou, Xinxiao Wu, Wentian Zhao, Jiebo Luo 0001, Yunde Jia
ICCV1
2018 Unsupervised Deep Learning of Mid-Level Video Representation for Action Recognition
abstract
Current deep learning methods for action recognition rely heavily on large scale labeled video datasets. Manually annotating video datasets is laborious and may introduce unexpected bias to train complex deep models for learning video representation. In this paper, we propose an unsupervised deep learning method which employs unlabeled local spatial-temporal volumes extracted from action videos to learn midlevel video representation for action recognition. Specifically, our method simultaneously discovers mid-level semantic concepts by discriminative clustering and optimizes local spatial-temporal features by two relatively small and simple deep neural networks. The clustering generates semantic visual concepts that guide the training of the deep networks, and the networks in turn guarantee the robustness of the semantic concepts. Experiments on the HMDB51 and the UCF101 datasets demonstrate the superiority of the proposed method, even over several supervised learning methods.
Jingyi Hou, Xinxiao Wu, Jin Chen 0009, Jiebo Luo 0001, Yunde Jia
AAAI1
2018 A discriminative structural model for joint segmentation and recognition of human actions
Cuiwei Liu, Jingyi Hou, Xinxiao Wu, Yunde Jia
Multim. Tools Appl.2
2018 Content-Attention Representation by Factorized Action-Scene Network for Action Recognition
abstract
During action recognition in videos, irrelevant motions in the background can greatly degrade the performance of recognizing specific actions with which we actually concern ourself here. In this paper, a novel deep neural network, called factorized action-scene network (FASNet), is proposed to encode and fuse the most relevant and informative semantic cues for action recognition. Specifically, we decompose the FASNet into two components. One is a newly designed encoding network, named content attention network (CANet), which encodes local spatial-temporal features to learn the action representations with good robustness to the noise of irrelevant motions. The other is a fusion network, which integrates the pretrained CANet to fuse the encoded spatial-temporal features with contextual scene feature extracted from the same video, for learning more descriptive and discriminative action representations. Moreover, different from the existing deep learning based tasks for generic action recognition, which applies softmax loss function as the training guidance, we formulate two loss functions for guiding the proposed model to accomplish more specific action recognition tasks, i.e., the multilabel correlation loss for multilabel action recognition and the triplet loss for complex event detection. Extensive experiments on the Hollywood2 dataset and the TRECVID MEDTest 14 dataset show that our method achieves superior performance compared with the state-of-the-art methods.
Jingyi Hou, Xinxiao Wu, Yuchao Sun, Yunde Jia
IEEE Trans. Multim.1
2016 Multimedia event detection via deep spatial-temporal neural networks
abstract
This paper proposes a novel method using deep spatial-temporal neural networks based on deep Convolutional Neural Network (CNN) for multimedia event detection. To sufficiently take advantage of the motion and appearance information of events from videos, our networks contain two branches: a temporal neural network and a spatial neural network. The temporal neural network captures motion information by Recurrent Neural Networks with the mutation of gated recurrent unit. The spatial neural network catches object information by using the deep CNN, to encode the CNN features as a bag of semantics with more discriminative representations. Both the temporal and spatial features are beneficial for event detection in a fully coupled way. Finally, we employ the generalized multiple kernel learning method to effectively fuse these two types of heterogeneous and complementary features for action recognition. Experiments on TRECVID MEDTest 14 dataset show that our method achieves better performance than the state of the art.
Jingyi Hou, Xinxiao Wu, Feiwu Yu, Yunde Jia
ICME1