VLDB 2026 Research / reviewers in the wild / expert
Dejie Yang
dblp:251/0232
· DBLP profile ↗
10ranked-venue papers
7as first author
8since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 4 first-author · 5 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FedAligns: A Federated Learning-Based IoT Intrusion Detection Method With Direction-Aligned UpdatesabstractIoT intrusion detection faces critical challenges from device heterogeneity, non-iid data distributions, and malicious updates. To address these, we propose FedAligns, a robust and efficient federated learning intrusion detection framework. First, we design a Temporal Attention Autoencoder that integrates TCN, GRU, and self-attention to capture local patterns and global dependencies in complex IoT traffic, enhancing feature representation. Second, we introduce a dual-aggregation algorithm, AlignMSE, to tackle non-iid heterogeneity and malicious threats. Its first phase filters anomalous updates via temporal direction alignment and sign consistency to ensure robust aggregation. Its second phase dynamically allocates weights based on local Mean Squared Error, prioritizing high-performing clients to better adapt to non-iid data. Experiments on N-BaIoT and ToN IoT datasets show FedAligns significantly outperforms traditional methods in accuracy and robustness, demonstrating superior resilience against malicious updates in non-iid scenarios. Keyuan Qiu, Zhejie Xu, Jianlin Lu, Xiaohui Xu, Shuangshi Zhao, Meifang Yan, Tao Luo 0016, Dejie Yang, Zhigang Li 0004 |
IEEE Internet Things J. | 8 |
| 2026 | MASC: Joint motion-aware and VLM-semantic augmentation for contextual traffic anomaly detection
Chengcheng Xu 0001, Dejie Yang, Qi Ai, Haiguang Lai |
Pattern Recognit. | 3 |
| 2025 | PlanLLM: Video Procedure Planning with Refinable Large Language ModelsabstractVideo procedure planning, i.e., planning a sequence of action steps given the video frames of start and goal states, is an essential ability for embodied AI. Recent works utilize Large Language Models (LLMs) to generate enriched action step description texts to guide action step decoding. Although LLMs are introduced these methods decode the action steps into a closed-set of one-hot vectors, limiting the model's capability of generalizing to new steps or tasks. Additionally, fixed action step descriptions based on world-level commonsense may contain noise in specific instances of visual states. In this paper, we propose PlanLLM, a cross-modal joint learning framework with LLMs for video procedure planning. We propose an LLM-Enhanced Planning module which fully uses the generalization ability of LLMs to produce free-form planning output and to enhance action step decoding. We also propose Mutual Information Maximization module to connect world-level commonsense of step descriptions and sample-specific information of visual states, enabling LLMs to employ the reasoning ability to generate step sequences. With the assistance of LLMs, our method can both closed-set and open vocabulary procedure planning tasks. Our PlanLLM achieves superior performance on three benchmarks, demonstrating the effectiveness of our designs. Dejie Yang, Zijing Zhao 0001, Yang Liu 0105 |
AAAI | 1 |
| 2025 | AR-VRM: Imitating Human Motions for Visual Robot Manipulation with Analogical ReasoningabstractVisual Robot Manipulation (VRM) aims to enable a robot to follow natural language instructions based on robot states and visual observations, and therefore requires costly multi-modal data. To compensate for the deficiency of robot data, existing approaches have employed vision-language pretraining with large-scale data. However, they either utilize web data that differs from robotic tasks, or train the model in an implicit way (e.g., predicting future frames at the pixel level), thus showing limited generalization ability under insufficient robot data. In this paper, we propose to learn from large-scale human action video datasets in an explicit way (i.e., imitating human actions from hand keypoints), introducing Visual Robot Manipulation with Analogical Reasoning (AR-VRM). To acquire action knowledge explicitly from human action videos, we propose a keypoint Vision-Language Model (VLM) pretraining scheme, enabling the VLM to learn human action knowledge and directly predict human hand keypoints. During fine-tuning on robot data, to facilitate the robotic arm in imitating the action patterns of human motions, we first retrieve human action videos that perform similar manipulation tasks and have similar historical observations , and then learn the Analogical Reasoning (AR) map between human hand keypoints and robot components. Taking advantage of focusing on action keypoints instead of irrelevant visual cues, our method achieves leading performance on the CALVIN benchmark {and real-world experiments}. In few-shot scenarios, our AR-VRM outperforms previous methods by large margins , underscoring the effectiveness of explicitly imitating human actions under data scarcity. Dejie Yang, Zijing Zhao 0004, Yang Liu 0105 |
ICCV | 1 |
| 2025 | Hierarchical Sub-action Tree for Continuous Sign Language RecognitionabstractContinuous sign language recognition (CSLR) aims to transcribe untrimmed videos into glosses, which are typically textual words. Recent studies indicate that the lack of large datasets and precise annotations has become a bottleneck for CSLR due to insufficient training data. To address this, some works have developed cross-modal solutions to align visual and textual modalities. However, they typically extract textual features from glosses without fully utilizing their knowledge. In this paper, we propose the Hierarchical Sub-action Tree (HST), termed HST-CSLR, to efficiently combine gloss knowledge with visual representation learning. By incorporating gloss-specific knowledge from large language models, our approach leverages textual information more effectively. Specifically, we construct an HST for textual information representation, aligning visual and textual modalities step-by-step and benefiting from the tree structure to reduce computational complexity. Additionally, we impose a contrastive alignment enhancement to bridge the gap between the two modalities. Experiments on four datasets (PHOENIX-2014, PHOENIX-2014T, CSL-Daily, and Sign Language Gesture) demonstrate the effectiveness of our HST-CSLR. Code and model are available at: https://github.com/Federfallt/HST-CSLR.git. Dejie Yang, Xinjie Gao |
ICME | 1 |
| 2024 | Active Object Detection with Knowledge Aggregation and Distillation from Large ModelsabstractAccurately detecting active objects undergoing state changes is essential for comprehending human interactions and facilitating decision-making. The existing methods for active object detection (AOD) primarily rely on visual appearance of the objects within input, such as changes in size, shape and relationship with hands. However, these visual changes can be subtle, posing challenges, particularly in scenarios with multiple distracting no-change instances of the same category. We observe that the state changes are often the result of an interaction being performed upon the object, thus propose to use informed priors about object related plausible interactions (including semantics and visual appearance) to provide more reliable cues for AOD. Specifically, we propose a knowledge aggregation procedure to integrate the aforementioned informed priors into oracle queries within the teacher decoder, offering more object affordance commonsense to locate the active object. To streamline the inference process and reduce extra knowledge inputs, we propose a knowledge distillation approach that encourages the student decoder to mimic the detection capabilities of the teacher decoder using the oracle query by replicating its predictions and attention. Our proposed framework achieves state-of-the-art performance on four datasets, namely Ego4D, Epic-Kitchens, MECCANO, and 100DOH, which demonstrates the effectiveness of our approach in improving AOD. The code and models are available at https://github.com/idejie/KAD.git. Dejie Yang, Yang Liu 0105 |
CVPR | 1 |
| 2024 | 3D Vision and Language Pretraining with Large-Scale Synthetic Data
Dejie Yang, Wentao Mo, Qingchao Chen, Siyuan Huang 0001, Yang Liu 0105 |
IJCAI | 1 |
| 2023 | Recent Advances in Class-Incremental Learning
Dejie Yang, Minghang Zheng, Weishuai Wang, Yang Liu 0105 |
ICIG (2) | 1 |
| 2020 | Deep Semantic-Alignment Hashing for Unsupervised Cross-Modal RetrievalabstractDeep hashing methods have achieved tremendous success in cross-modal retrieval, due to its low storage consumption and fast retrieval speed. In real cross-modal retrieval applications, it's hard to obtain label information. Recently, increasing attention has been paid to unsupervised cross-modal hashing. However, existing methods fail to exploit the intrinsic connections between images and their corresponding descriptions or tags (text modality). In this paper, we propose a novel Deep Semantic-Alignment Hashing (DSAH) for unsupervised cross-modal retrieval, which sufficiently utilizes the co-occurred image-text pairs. DSAH explores the similarity information of different modalities and we elaborately design a semantic-alignment loss function, which elegantly aligns the similarities between features with those between hash codes. Moreover, to further bridge the modality gap, we innovatively propose to reconstruct features of one modality with hash codes of the other one. Extensive experiments on three cross-modal retrieval datasets demonstrate that DSAH achieves the state-of-the-art performance. Dejie Yang, Dayan Wu, Wanqian Zhang, Haisu Zhang, Bo Li 0063, Weiping Wang 0005 |
ICMR | 1 |
| 2019 | Network Clustering Analysis Using Mixture Exponential-Family Random Graph Models and Its Application in Genetic Interaction DataabstractMOTIVATION: Epistatic miniarrary profile (EMAP) studies have enabled the mapping of large-scale genetic interaction networks and generated large amounts of data in model organisms. It provides an incredible set of molecular tools and advanced technologies that should be efficiently understanding the relationship between the genotypes and phenotypes of individuals. However, the network information gained from EMAP cannot be fully exploited using the traditional statistical network models. Because the genetic network is always heterogeneous, for example, the network structure features for one subset of nodes are different from those of the left nodes. Exponential-family random graph models (ERGMs) are a family of statistical models, which provide a principled and flexible way to describe the structural features (e.g., the density, centrality, and assortativity) of an observed network. However, the single ERGM is not enough to capture this heterogeneity of networks. In this paper, we consider a mixture ERGM (MixtureEGRM) networks, which model a network with several communities, where each community is described by a single EGRM. RESULTS: EM algorithm is a classical method to solve the mixture problem, however, it will be very slow when the data size is huge in the numerous applications. We adopt an efficient novel online graph clustering algorithm to classify the graph nodes and estimate the ERGM parameters for the MixtureERGM. In comparison studies, the MixtureERGM outperforms the role analysis for the network cluster in which the mixture of exponential-family random graph model is developed for many ego-network according to their roles. One genetic interaction network of yeast and two real social networks (provided as supplemental materials, which can be found on the Computer Society Digital Library at http://doi.ieeecomputersociety.org/10.1109/TCBB.2017.2743711) show the wide potential application of the MixtureERGM. Yishu Wang 0003, Huaying Fang, Dejie Yang, Hongyu Zhao 0003, Minghua Deng |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |