EDBT 2026 Demo / reviewers in the wild / expert
Jiale Yu
dblp:121/3488
· DBLP profile ↗
17ranked-venue papers
6as first author
12since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 8 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Boosting Few-Shot Human-Object Recognition With Hierarchical Correlation LearningabstractIdentifying novel human-object interaction (HOI) classes with scarce data is a challenging and crucial task in computer vision. Existing methods mainly use coarse global visual information to build class prototypes in meta-learning. Despite their promising results, these methods often fail to capture fine-grained interaction semantics and effectively learn from data with low inter-class variance, leading to suboptimal performance in distinguishing similar categories. To overcome these issues, we propose a new model called hierarchical relation network for few-shot HOI recognition (FS-HOI). This model integrates multi-level interaction clues, spanning from coarse to fine-grained, to enhance HOI features. It employs a unified graph network to capture intra- and inter-relationships among human parts with contextual information, augmented by language-guided attention for semantic mining within each interactive sub-graph. In contrast to conventional methods that depend on global class prototype comparisons, our approach advances metric learning by integrating contrastive mechanisms, utilizing rich instance pairs as comparative references to effectively address inter-class variance. Furthermore, a graph relation network leverages prior knowledge of unknown HOIs, embedding task-specific features into contrastive instances. Our method establishes a new state-of-the-art on three few-shot HOI datasets, with substantial performance gains and ablation studies confirming the efficacy of each component. Jiale Yu, Baopeng Zhang, Zhu Teng, Jianping Fan 0007 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Local Causal Discovery Without Causal SufficiencyabstractLocal causal discovery is crucial for revealing the causal relationships between specific variables from data. Existing local causal discovery algorithms are designed under the assumption of causal sufficiency, which states that there are no latent common causes for two or more of the observed variables in data. However, the assumption of causal sufficiency is often violated in practice. To address this issue, we first propose the local Maximal Ancestral Graph (MAG), referred to as LocalMAG, to describe the local causal relationships of the target variable in the MAG. Then, we propose a local causal discovery algorithm without the assumption of causal sufficiency, called LatentLCD, to learn the LocalMAG. Specifically, LatentLCD first uses the traditional parents and children discovery algorithm to identify the local causal skeleton that includes latent variables and verifies it theoretically. It then identifies bidirectional edges by determining whether both the target variable and its adjacent variables are colliders, thereby identifying latent variables in the local structure of the target variable. Extensive experiments on synthetic datasets have validated that the proposed LatentLCD algorithm significantly outperforms the state-of-the-art methods. Zhaolong Ling, Jiale Yu, Yiwen Zhang 0001, Debo Cheng, Peng Zhou 0006, Bingbing Jiang 0001, Kui Yu |
AAAI | 2 |
| 2025 | GROVE: A Generalized Reward for Learning Open-Vocabulary Physical SkillabstractLearning open-vocabulary physical skills for simulated agents presents a significant challenge in Artificial Intelligence (AI). Current Reinforcement Learning (RL) approaches face critical limitations: manually designed rewards lack scalability across diverse tasks, while demonstration-based methods struggle to generalize beyond their training distribution. We introduce GROVE, a generalized reward framework that enables open-vocabulary physical skill learning without manual engineering or task-specific demonstrations. Our key insight is that Large Language Models (LLMs) and Vision Language Models (VLMs) provide complementary guidance—LLMs generate precise physical constraints capturing task requirements, while VLMs evaluate motion semantics and naturalness. Through an iterative design process, VLM-based feedback continuously refines LLM-generated constraints, creating a self-improving reward system. To bridge the domain gap between simulation and natural images, we develop Pose2CLIP, a lightweight mapper that efficiently projects agent poses directly into semantic feature space without computationally expensive rendering. Extensive experiments across diverse embodiments and learning paradigms demonstrate GROVE’s effectiveness, achieving 22.2% higher motion naturalness and 25.7% better task completion scores while training 8.4× faster than previous methods. These results establish a new foundation for scalable physical skill acquisition in simulated environments. Jieming Cui, Tengyu Liu, Jiale Yu, Ran Song 0001, Wei Zhang 0021, Yixin Zhu 0001, Siyuan Huang 0001 |
CVPR | 4 |
| 2025 | OV-DAVEL: Towards Open-Vocabulary Dense Audio-Visual Event Localization in Untrimmed VideosabstractThe Dense Audio-Visual Event Localization (DAVEL) task aims to temporally localize all audio-visual events within untrimmed videos. Existing methods typically operate under a closed-set assumption, which limits their ability to generalize to test videos containing previously unseen event categories-an essential capability for open-world scenarios. To this end, we propose a novel task setting, Open-Vocabulary Dense Audio-Visual Event Localization (OV-DAVEL), along with a one-stage method Open-DAVTR designed to detect events that were not observed during training. Open-DAVTR consists of two core components: a class-agnostic foreground-aware generator and a multi-modal semantic-aware classifier. Specifically, the generator is a Detection Transformer-based module that produces event proposals while adaptively attending to discriminative foreground snippets for downstream classification. The classifier leverages rich temporal representations and context-aware textual semantics to effectively recognize events, regardless of whether they were seen during training. In addition, we establish comprehensive OV-DAVEL benchmarks across various settings. Experimental results show that our model significantly outperforms existing baselines in detecting both seen and unseen events, highlighting its effectiveness in open-vocabulary event localization. Code and data are available at: https://github.com/yujialele/OV-DEAVEL. Jiale Yu, Baopeng Zhang, Zhu Teng, Jianping Fan 0007 |
ACM Multimedia | 1 |
| 2025 | ACR: Adaptive Computation Reuse for Video Analytics in Collaborative Edge ComputingabstractVideo analytics typically requires substantial computational resources and energy. In edge computing application scenarios (e.g., smart cities), multiple users and devices may be spatially proximate, leading to offloaded tasks with high similarity and redundant computation. Computational results from previously executed tasks can be cached and reused for subsequent tasks based on similarity to enhance system efficiency. However, existing computation reuse methods often employ fixed similarity thresholds tailored to specific tasks, struggling to adapt to dynamically changing scenarios. This results in low reuse accuracy under low similarity thresholds and high latency under high similarity thresholds. Additionally, while edge servers possess larger storage capacities, they significantly increase real-time retrieval overhead. To address these issues, this paper proposes ACR, an adaptive computation reuse framework based on edge-cloud collaboration. ACR leverages the localized deployment of edge gateways to reduce cache query latency and introduces a deep Q-Network (DQN) algorithm with n-step temporal difference (TD) to adaptively adjust similarity thresholds. Our evaluation results demonstrate that, on typical urban surveillance datasets, ACR effectively balances overall system latency and reuse accuracy compared to other computation reuse methods. Jiale Yu, Minghua Zhu |
SMC | 1 |
| 2024 | Education From Short Video: A Novel Educational Pattern Inspired by Entertainment VideosabstractBlended learning has gained popularity for its flexibility and technological integration. However, many current models simply replicate offline classrooms in virtual environments, lacking engagement for younger students. While gamified education fosters student motivation, it fails to enhance deep knowledge comprehension or assist in automating assessment processes. This paper introduces a novel educational approach, Education From Short Video (EFSV), which uniquely advocates for ‘Short-Video-Izing’ educational content and automating student assessments. We propose a model that transforms content into interactive short videos by blending AI-driven and manual techniques, moving away from traditional definition-based methods. Additionally, a Short Video Platform of Education (SVPE) is developed, where students act as both content creators and viewers. On this platform, student interactions like ‘comments’ and ‘likes’ automate final grading, significantly reducing the workload of teachers. Experimental results demonstrate that this model enhances students’ learning effectiveness and enjoyment. Jiale Yu, Zhiwei Zheng |
ISPA | 1 |
| 2024 | OpenAVE: Moving towards Open Set Audio-Visual Event LocalizationabstractAudio-Visual Event (AVE) Localization aims to identify and classify video segments that are both audible and visible, a field that has seen substantial progress in recent years. Existing methods operate under a closed-set assumption and struggle to recognize unknown events in open-world scenarios. To better adapt to real-life applications, we introduce the Open Set Audio-Visual Event Localization task and propose a novel and effective network called OpenAVE based on evidential deep learning. To the best of our knowledge, this is the first effort to address this challenge. Our approach encompasses deep evidential AVE classification and event-relevant prediction, targeting the nuanced demands of open-set environments. The deep evidential AVE classification manages event classification uncertainty by extracting class evidence from segment-specific representations enriched with multi-scale context. To effectively distinguish between unknown events and background segments, event-relevant prediction utilizes positive-unlabeled learning. Futhermore, a learnable Gaussian-prior prediction branch is adopted to enhance the performance of event-relevant prediction. Experimental results demonstrate that OpenAVE significantly outperforms state-of-the-art models on the Audio-Visual Event dataset, confirming the effectiveness of our proposed method. Jiale Yu, Baopeng Zhang, Zhu Teng, Jianping Fan 0007 |
ACM Multimedia | 1 |
| 2023 | Hierarchical Reasoning Network with Contrastive Learning for Few-Shot Human-Object Interaction RecognitionabstractFew-shot learning (FSL) for human-object interaction aims at classifying samples of new unseen HOI classes with only a few labeled samples available. Although progress has been made in few-shot human-object interaction, most of the existing methods encounter two issues in handling fine-grained interactions: the inability to capture more subtle interactive clues and the inadequacy in learning from data with low inter-class variance. To tackle the first issue, we propose a hierarchical reasoning network to integrate multi-level interactive clues (from coarse to fine-grained) for strengthening HOI representations. The hierarchical relation module mainly captures and aggregates more discriminative relation information among human parts at multiple levels (including the human instance, action region, and body part levels) and objects via a unified graph and exploits a language-guided attentive fusion way to highlight informative features of each interaction level. To address the second issue, we introduce a contrastive learning mechanism to alleviate the inter-class variance. Compared with the previous ProtoNet-based methods, our model generates more discriminative representations for low inter-class variance data, since it makes full use of potential contrastive pairs in each training episode. Extensive experimental results on two standard benchmarks demonstrate that the proposed model performs favorably against state-of-the-art FS-HOI methods. Jiale Yu, Baopeng Zhang, Zhu Teng |
ACM Multimedia | 1 |
| 2023 | Self-supervised group meiosis contrastive learning for EEG-based emotion recognition
Haoning Kan, Jiale Yu, Jiajin Huang, Heqian Wang |
Appl. Intell. | 2 |
| 2022 | PiLSL: pairwise interaction learning-based graph neural network for synthetic lethality prediction in human cancersabstractMOTIVATION: Synthetic lethality (SL) is a type of genetic interaction in which the simultaneous inactivation of two genes leads to cell death, while the inactivation of a single gene does not affect the cell viability. It can effectively expand the range of anti-cancer therapeutic targets. SL interactions are identified mainly by experimental screening and computational prediction. Recent machine-learning methods mostly learn the representation of each gene individually, ignoring the representation of the pairwise interaction between two genes. In addition, the mechanisms of SL, the key to translating SL into cancer therapeutics, are often unclear. RESULTS: To fill the gaps, we propose a pairwise interaction learning-based graph neural network (GNN) named PiLSL to learn the representation of pairwise interaction between two genes for SL prediction. First, we construct an enclosing graph for each pair of genes from a knowledge graph. Secondly, we design an attentive embedding propagation layer in a GNN to discriminate the importance among the edges in the enclosing graph and to learn the latent features of the pairwise interaction from the weighted enclosing graph. Finally, we further fuse the latent features with explicit features extracted from multi-omics data to obtain powerful gene representations for SL prediction. Extensive experimental results demonstrate that PiLSL outperforms the best baseline by a large margin and generalizes well under three realistic scenarios. Besides, PiLSL provides an explanation of SL mechanisms via the weighted paths in the enclosing graphs by attention mechanism. AVAILABILITY AND IMPLEMENTATION: Our source code is available at https://github.com/JieZheng-ShanghaiTech/PiLSL. Xin Liu 0027, Jiale Yu, Beiyuan Yang, Shike Wang, Fang Bai, Jie Zheng 0002 |
Bioinform. | 2 |
| 2022 | A deep learning-based method for pixel-level crack detection on concrete bridgesabstractAbstract Crack detection of the concrete bridge is an essential index for the safety assessment of bridge structure. It is more important to check the whole structure than to check the accuracy in the damage assessment. However, the traditional deep learning model method cannot completely detect the crack structure, which challenges image‐based crack detection. For this reason, we propose deep bridge crack classification (DBCC)‐Net as a classification‐based deep learning network. By pruning the Yolox, the regression problem of the target detection is converted to the binary classification problem to avoid the network performance degradation caused by the translation invariance of the convolutional neural network (CNN). In addition, the network post‐processing and a two‐stage crack detection strategy are proposed to enable the network to detect cracks and extract crack morphology in high‐resolution images quickly. In the first stage, DBCC‐Net realizes the coarse extraction of crack position based on image slice classification. In the second stage, the complete crack morphology is extracted from the location suggested by the semantic segmentation network. Experimental results show that the proposed two‐stage method has 19 frames per second (FPS) and 0.79 Miou (mean intersection over union) at the actual bridge images with 2560×2560 pixels. Although FPS is reduced, the Miou value is 7.8% higher than other methods, proving this paper's practical value. Ji Kun, Jiale Yu |
IET Image Process. | 3 |
| 2021 | An Improved Deep Relation Network for Action Recognition in Still ImagesabstractContextual information has been widely utilized in visual recognition tasks. This is especially true for action recognition, because contextual information such as objects interacting with human and the scene where the action is performed is inseparable from action categories. To this end, we propose an efficient relation module that combines Human-Object and Scene-Object relations for action recognition. Specifically, Human-Object interaction submodule can capture more accurate appearance and spatial relation to build human-object interaction pairs. And Scene- Object interaction submodule can learn the probability of the objects involved in the scene to help discover the key interaction pair. We conduct extensive experiments on Stanford 40 and Pascal Voc 2012 Action datasets to verify our model, and experimental results show that our method achieves superior performance on these two datasets. Especially, we gain the best results on the Stanford 40 dataset compared with state-of-the-arts. Wei Wu 0032, Jiale Yu |
ICASSP | 2 |
| 2020 | A Part Fusion Model for Action Recognition in Still Images
Wei Wu 0032, Jiale Yu |
ICONIP (1) | 2 |
| 2020 | An Improved Bilinear Pooling Method for Image-Based Action RecognitionabstractAction recognition in still images is a challenging task because of the complexity of human motions and the variance of background in the same action category. And some actions typically occur in fine-grained categories, with little visual differences between these categories. So extracting discriminative features or modeling various semantic parts is essential for image-based action recognition. Many methods apply expensive manual annotations to learn discriminative parts information for action recognition, which may severely hinder potential applications in real life. In recent years, bilinear pooling method has shown its effectiveness for image classification due to its learning distinctive features automatically. Inspired by this model, in this paper, an improved bilinear pooling method is proposed for image-based action recognition. The previous bilinear pooling approaches contain lots of noisy background or harmful feature information, which limit their application for action recognition. In our method, the attention mechanism is introduced into hierarchical bilinear pooling framework with mask aggregation. The proposed model can generate the distinctive and RoI-aware feature information by combining multiple attention mask maps from the channel and spatial-wise attention features. Specifically, our method makes the network to pay more attention to discriminative region of the vital objects in an image. We verify our model on the two challenging datasets: 1) Stanford 40 action dataset and 2) our action dataset that includes 60 categories. Experimental results demonstrate the effectiveness of our approach, which is superior to the traditional and state-of-the-art methods. Wei Wu 0032, Jiale Yu |
ICPR | 2 |
| 2020 | Multi-hop Reading Comprehension across Documents with Path-based Graph Convolutional NetworkabstractMulti-hop reading comprehension across multiple documents attracts much attentions recently. In this paper, we propose a novel approach to tackle this multi-hop reading comprehension problem. Inspired by the human reasoning processing, we introduce a path-based graph with reasoning paths which extracted from supporting documents. The path-based graph can combine both the idea of the graph-based and path-based approaches, so it is better for multi-hop reasoning. Meanwhile, we propose Gated-GCN to accumulate evidences on the path-based graph, which contains a new question-aware gating mechanism to regulate the usefulness of information propagating across documents and add question information during reasoning. We evaluate our approach on WikiHop dataset, and our approach achieves the the-state-of-art accuracy against previous published approaches. Especially, our ensemble model surpasses the human performance by 4.2%. Zeyun Tang, Yongliang Shen 0001, Xinyin Ma, Jiale Yu, Weiming Lu 0001 |
IJCAI | 5 |
| 2019 | Concept Extraction and Prerequisite Relation Learning from Educational DataabstractPrerequisite relations among concepts are crucial for educational applications. However, it is difficult to automatically extract domain-specific concepts and learn the prerequisite relations among them without labeled data.In this paper, we first extract high-quality phrases from a set of educational data, and identify the domain-specific concepts by a graph based ranking method. Then, we propose an iterative prerequisite relation learning framework, called iPRL, which combines a learning based model and recovery based model to leverage both concept pair features and dependencies among learning materials. In experiments, we evaluated our approach on two real-world datasets Textbook Dataset and MOOC Dataset, and validated that our approach can achieve better performance than existing methods. Finally, we also illustrate some examples of our approach. Weiming Lu 0001, Jiale Yu, Chenhao Jia |
AAAI | 3 |
| 2019 | Metro maps for efficient knowledge learning by summarizing massive electronic textbooks
Weiming Lu 0001, Pengkun Ma, Jiale Yu, Baogang Wei |
Int. J. Document Anal. Recognit. | 3 |