EDBT 2026 Demo / reviewers in the wild / expert
Jingyi Hou
dblp:173/4937
· DBLP profile ↗
12ranked-venue papers
7as first author
5since 2021 · last 2025
0000-0003-0172-3115ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 6 first-authorArtificial intelligence and machine learning · 5 · 3 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Video understanding and tracking · 36% Knowledge representation and reasoning · 18% Vision and language · 16% |
Topics — the 15 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Video understanding and tracking
action recognition |
1.1 | 3 | 2020 | Confidence-Guided Self Refinement for Action Prediction in Untrimmed Videos · IEEE Trans. Image Process. 2020 Content-Attention Representation by Factorized Action-Scene Network for Action Recognition · IEEE Trans. Multim. 2018 Unsupervised Deep Learning of Mid-Level Video Representation for Action Recognition · AAAI 2018 |
Computer vision › Video understanding and tracking
action anticipation |
0.9 | 2 | 2021 | Spatial-Temporal Relation Reasoning for Action Prediction in Videos · Int. J. Comput. Vis. 2021 Confidence-Guided Self Refinement for Action Prediction in Untrimmed Videos · IEEE Trans. Image Process. 2020 |
Computer vision › Vision and language
video captioning |
0.8 | 2 | 2020 | Joint Commonsense and Relation Reasoning for Image and Video Captioning · AAAI 2020 Joint Syntax Representation Learning and Visual Cue Translation for Video Captioning · ICCV 2019 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
latent factor model |
0.8 | 1 | 2024 | Discovering Predictable Latent Factors for Time Series Forecasting · IEEE Trans. Knowl. Data Eng. 2024 |
Machine learning › Time series and sequential data › time series analysis
time series forecasting |
0.8 | 1 | 2024 | Discovering Predictable Latent Factors for Time Series Forecasting · IEEE Trans. Knowl. Data Eng. 2024 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › temporal reasoning
spatio-temporal reasoning |
0.5 | 1 | 2021 | Spatial-Temporal Relation Reasoning for Action Prediction in Videos · Int. J. Comput. Vis. 2021 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
commonsense reasoning |
0.4 | 1 | 2020 | Joint Commonsense and Relation Reasoning for Image and Video Captioning · AAAI 2020 |
Computer vision › Vision and language
image captioning |
0.4 | 1 | 2020 | Joint Commonsense and Relation Reasoning for Image and Video Captioning · AAAI 2020 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
relational reasoning |
0.4 | 1 | 2020 | Joint Commonsense and Relation Reasoning for Image and Video Captioning · AAAI 2020 |
Computer vision › Video understanding and tracking › action recognition
spatio-temporal action recognition |
0.4 | 1 | 2020 | Confidence-Guided Self Refinement for Action Prediction in Untrimmed Videos · IEEE Trans. Image Process. 2020 |
Computer vision › Video understanding and tracking › event recognition
complex event detection |
0.3 | 1 | 2018 | Content-Attention Representation by Factorized Action-Scene Network for Action Recognition · IEEE Trans. Multim. 2018 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
discriminative clustering |
0.3 | 1 | 2018 | Unsupervised Deep Learning of Mid-Level Video Representation for Action Recognition · AAAI 2018 |
Machine learning › Learning paradigms › unsupervised learning
unsupervised video representation learning |
0.3 | 1 | 2018 | Unsupervised Deep Learning of Mid-Level Video Representation for Action Recognition · AAAI 2018 |
Machine learning › Deep learning architectures and training
attention mechanism |
0.1 | 1 | 2018 | Content-Attention Representation by Factorized Action-Scene Network for Action Recognition · IEEE Trans. Multim. 2018 |
Machine learning › Deep learning architectures and training › attention mechanism
content attention |
0.1 | 1 | 2018 | Content-Attention Representation by Factorized Action-Scene Network for Action Recognition · IEEE Trans. Multim. 2018 |
Methods — techniques the papers use, named apart from their topics
transformer · 0.8deep latent dynamics models · 0.8relation reasoning · 0.5sparse self-attention · 0.4semantic graph · 0.4self-refining network · 0.4iterative learning · 0.4gumbel-softmax · 0.4confidence learning · 0.4commonsense reasoning · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Image sentiment analysis based on distillation and sentiment region localization networkabstractAbstract Accurately identifying the emotions in images is crucial for sentiment content analysis. To detect local sentiment regions and acquire discriminative sentiment features, we propose a novel model named Distillation-guided and Contrastive-enhanced Sentiment Region Localization Network (DC-SRLN) to effectively complete image sentiment analysis. Two smart but heterogeneous SRLNs are designed first to pursue local sentiment regions. Then an innovative contrastive learning mode is implemented between global and local features to further enhance the discriminative ability of the sentiment features. Third, the enhanced global and local sentiment features are seamlessly integrated to guide each SRLN accurately capture local sentiment regions. Finally, an adaptive feature fusion module is created to fuse the heterogeneous features from the two SRLNs and generate a new multi-view multi-granularity sentiment semantics with more discriminative ability for image sentiment analysis. Extensive experimental results on three prevailing datasets, namely Twitter I, FI, and ArtPhoto, exhibit that DC-SRLN achieves satisfactory accuracies of 93.2%, 80.6%, and 78.7%, respectively, outperforming recent state-of-the-art baselines. Moreover, DC-SRLN needs less training time, demonstrating its high practicality. The code of DC-SRLN is freely available at https://github.com/Riley6868/DC-SRLN. Hongbin Zhang 0004, Ya Feng, Jingyi Hou, Guangli Li |
Comput. J. | 4 |
| 2024 | Type-adaptive graph Transformer for heterogeneous information networks
Yanzhe Huang, Jingyi Hou, Zhijie Liu 0001 |
Appl. Intell. | 3 |
| 2024 | Discovering Predictable Latent Factors for Time Series ForecastingabstractModern temporal modeling methods, such as Transformer and its variants, have demonstrated remarkable capabilities in handling sequential data from specific domains like language and vision. Though achieving high performance with large-scale data, they often have redundant or unexplainable structures. When encountering some real-world datasets with limited observable variables that can be affected by many unknown factors, these methods may struggle to identify meaningful patterns and dependencies inherent in data, and thus, the modeling becomes unstable and unpredictable. To tackle this critical issue, in this article, we develop a novel algorithmic framework for inferring latent factors implied by the observed temporal data. The inferred factors are used to form multiple predictable and independent signal components that enable not only the reconstruction of future time series for accurate prediction but also sparse relation reasoning for long-term efficiency. To achieve this, we introduce three characteristics, i.e., predictability, sufficiency, and identifiability, and model these characteristics of latent factors via powerful deep latent dynamics models to infer the predictable signal components. Empirical results on multiple real datasets show the efficiency of our method for different kinds of time series forecasting tasks. Statistical analyses validate the predictability and interpretability of the learned latent factors. Jingyi Hou, Zhen Dong 0002, Zhijie Liu 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2023 | Edge AI as a Service: Configurable Model Deployment and Delay-Energy Optimization With Result Quality ConstraintsabstractThe breakthrough of artificial intelligence (AI) techniques has accelerated their applications in a wide range of industries, such as security protection, transportation, agriculture, and medical care. With the support of edge computing environments, providing latency guaranteed AI as a Service (AIaaS) can accelerate the deployment of data-intensive and computation-intensive AI applications and reduce the investment cost of the customers. However, the deployment architecture and working mechanism design, and performance optimization problems specific for AIaaS with configurable data quality and model complexity have not been studied in existing works. To address the problem, we propose a configurable model deployment architecture (CMDA) for edge AIaaS and present a flexible working mechanism by enabling the joint configuration of data quality ratios (DQRs) and model complexity ratios (MCRs) for the AI tasks. Along with commonly used resource allocation operations, the manager can improve the energy and delay performance of AI services with the desired quality of results (QoRs). We develop an energy-delay minimization problem under the framework of CMDA and propose a polynomial regression based relaxing method to solve the task configuration subproblem. We conduct experiments and simulations on the ImageNet classification and the common objects in context (COCO) object detection tasks using state-of-the-art deep learning models. We present the corresponding result quality tables (RQTs) and QoR regression models to illustrate the proposed method. The results of single task configuration and multi-task configuration and resource allocation on ImageNet classification and COCO object detection tasks demonstrate that the proposed method can achieve over$5\times$HDEC improvement compared with non-optimization schemes, and also show that joint configuration of DQR and MCR can achieve over$1.2\times$HDEC improvement compared with the methods that only configure DQR or MCR. Wenyu Zhang 0002, Sherali Zeadally, Wei Li 0074, Haijun Zhang 0001, Jingyi Hou, Victor C. M. Leung |
IEEE Trans. Cloud Comput. | 5 |
| 2021 | Spatial-Temporal Relation Reasoning for Action Prediction in Videos
Xinxiao Wu, Jingyi Hou, Hanxi Lin, Jiebo Luo 0001 |
Int. J. Comput. Vis. | 3 |
| 2020 | Joint Commonsense and Relation Reasoning for Image and Video CaptioningabstractExploiting relationships between objects for image and video captioning has received increasing attention. Most existing methods depend heavily on pre-trained detectors of objects and their relationships, and thus may not work well when facing detection challenges such as heavy occlusion, tiny-size objects, and long-tail classes. In this paper, we propose a joint commonsense and relation reasoning method that exploits prior knowledge for image and video captioning without relying on any detectors. The prior knowledge provides semantic correlations and constraints between objects, serving as guidance to build semantic graphs that summarize object relationships, some of which cannot be directly perceived from images or videos. Particularly, our method is implemented by an iterative learning algorithm that alternates between 1) commonsense reasoning for embedding visual regions into the semantic space to build a semantic graph and 2) relation reasoning for encoding semantic graphs to generate sentences. Experiments on several benchmark datasets validate the effectiveness of our prior knowledge-based approach. Jingyi Hou, Xinxiao Wu, Xiaoxun Zhang, Yayun Qi, Yunde Jia, Jiebo Luo 0001 |
AAAI | 1 |
| 2020 | Confidence-Guided Self Refinement for Action Prediction in Untrimmed VideosabstractMany existing methods formulate the action prediction task as recognizing early parts of actions in trimmed videos. In this paper, we focus on predicting actions from ongoing untrimmed videos where actions might not happen at the very beginning of videos. It is extremely challenging to predict actions in such untrimmed videos due to ambiguous or even no information of actions in the early parts of videos. To address this problem, we propose a prediction confidence that assesses the decision quality of a prediction model. Guided by the confidence, the model continuously refines the prediction results by itself with the increasing observed video frames. Specifically, we build a Self Prediction Refining Network (SPR-Net) which incrementally learns the confidence for action prediction. SPR-Net consists of three modules: a temporal hybrid network, an incremental confidence learner, and a self-refining Gumbel softmax sampler. The temporal hybrid network generates the action category distributions by integrating static scene and dynamic motion information. The incremental confidence learner calculates the confidence in an incremental manner, judging the extent to which the temporal hybrid network should believe its prediction result. The self-refining Gumbel softmax sampler models the mutual relationship between the prediction confidence and the category distribution, which enables them to be jointly learned in an end-to-end fashion. We also present a sparse self-attention mechanism to encode local spatio-temporal features into the frame-level motion representation to further improve the prediction performance. Extensive experiments on five datasets (i.e., UT-Interaction, BIT-Interaction, UCF101, THUMOS14, and ActivityNet) validate the effectiveness of the proposed method. Jingyi Hou, Xinxiao Wu, Jiebo Luo 0001, Yunde Jia |
IEEE Trans. Image Process. | 1 |
| 2019 | Joint Syntax Representation Learning and Visual Cue Translation for Video CaptioningabstractVideo captioning is a challenging task that involves not only visual perception but also syntax representation learning. Recent progress in video captioning has been achieved through visual perception, but syntax representation learning is still under-explored. We propose a novel video captioning approach that takes into account both visual perception and syntax representation learning to generate accurate descriptions of videos. Specifically, we use sentence templates composed of Part-of-Speech (POS) tags to represent the syntax structure of captions, and accordingly, syntax representation learning is performed by directly inferring POS tags from videos. The visual perception is implemented by a mixture model which translates visual cues into lexical words that are conditional on the learned syntactic structure of sentences. Thus, a video captioning task consists of two sub-tasks: video POS tagging and visual cue translation, which are jointly modeled and trained in an end-to-end fashion. Evaluations on three public benchmark datasets demonstrate that our proposed method achieves substantially better performance than the state-of-the-art methods, which validates the superiority of joint modeling of syntax representation learning and visual perception for video captioning. Jingyi Hou, Xinxiao Wu, Wentian Zhao, Jiebo Luo 0001, Yunde Jia |
ICCV | 1 |
| 2018 | Unsupervised Deep Learning of Mid-Level Video Representation for Action RecognitionabstractCurrent deep learning methods for action recognition rely heavily on large scale labeled video datasets. Manually annotating video datasets is laborious and may introduce unexpected bias to train complex deep models for learning video representation. In this paper, we propose an unsupervised deep learning method which employs unlabeled local spatial-temporal volumes extracted from action videos to learn midlevel video representation for action recognition. Specifically, our method simultaneously discovers mid-level semantic concepts by discriminative clustering and optimizes local spatial-temporal features by two relatively small and simple deep neural networks. The clustering generates semantic visual concepts that guide the training of the deep networks, and the networks in turn guarantee the robustness of the semantic concepts. Experiments on the HMDB51 and the UCF101 datasets demonstrate the superiority of the proposed method, even over several supervised learning methods. Jingyi Hou, Xinxiao Wu, Jin Chen 0009, Jiebo Luo 0001, Yunde Jia |
AAAI | 1 |
| 2018 | A discriminative structural model for joint segmentation and recognition of human actions
Cuiwei Liu, Jingyi Hou, Xinxiao Wu, Yunde Jia |
Multim. Tools Appl. | 2 |
| 2018 | Content-Attention Representation by Factorized Action-Scene Network for Action RecognitionabstractDuring action recognition in videos, irrelevant motions in the background can greatly degrade the performance of recognizing specific actions with which we actually concern ourself here. In this paper, a novel deep neural network, called factorized action-scene network (FASNet), is proposed to encode and fuse the most relevant and informative semantic cues for action recognition. Specifically, we decompose the FASNet into two components. One is a newly designed encoding network, named content attention network (CANet), which encodes local spatial-temporal features to learn the action representations with good robustness to the noise of irrelevant motions. The other is a fusion network, which integrates the pretrained CANet to fuse the encoded spatial-temporal features with contextual scene feature extracted from the same video, for learning more descriptive and discriminative action representations. Moreover, different from the existing deep learning based tasks for generic action recognition, which applies softmax loss function as the training guidance, we formulate two loss functions for guiding the proposed model to accomplish more specific action recognition tasks, i.e., the multilabel correlation loss for multilabel action recognition and the triplet loss for complex event detection. Extensive experiments on the Hollywood2 dataset and the TRECVID MEDTest 14 dataset show that our method achieves superior performance compared with the state-of-the-art methods. Jingyi Hou, Xinxiao Wu, Yuchao Sun, Yunde Jia |
IEEE Trans. Multim. | 1 |
| 2016 | Multimedia event detection via deep spatial-temporal neural networksabstractThis paper proposes a novel method using deep spatial-temporal neural networks based on deep Convolutional Neural Network (CNN) for multimedia event detection. To sufficiently take advantage of the motion and appearance information of events from videos, our networks contain two branches: a temporal neural network and a spatial neural network. The temporal neural network captures motion information by Recurrent Neural Networks with the mutation of gated recurrent unit. The spatial neural network catches object information by using the deep CNN, to encode the CNN features as a bag of semantics with more discriminative representations. Both the temporal and spatial features are beneficial for event detection in a fully coupled way. Finally, we employ the generalized multiple kernel learning method to effectively fuse these two types of heterogeneous and complementary features for action recognition. Experiments on TRECVID MEDTest 14 dataset show that our method achieves better performance than the state of the art. Jingyi Hou, Xinxiao Wu, Feiwu Yu, Yunde Jia |
ICME | 1 |