VLDB 2026 Research / reviewers in the wild / expert
Yingjie Xia
dblp:17/5533
· DBLP profile ↗
75ranked-venue papers
33as first author
32since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 23 · 9 first-author · 6 since 2021Artificial intelligence and machine learning · 20 · 10 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 5 first-author · 9 since 2021Computer networks · 9 · 4 first-author · 8 since 2021Databases, data management, data science and information retrieval · 6 · 4 first-authorSoftware engineering, systems software and programming languages · 5 · 3 since 2021Security and privacy · 3 · 2 since 2021Theory of computation · 3 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VerilogLAVD: LLM-Aided Pattern Generation for Verilog CWE DetectionabstractLLMs often fail in hardware vulnerability detection due to the intrinsic semantic concurrency of HDLs (Hardware Description Language), where vulnerabilities arise from the interaction of multiple concurrent execution statements rather than a single sequential execution path.Existing LLM-based methods struggle to capture the concurrency features.To address the problem, we propose VerilogLAVD, a LLM-Aided Vulnerability Detection framework by generating executable Traversal Detection Patterns (TDPs), i.e. the rules describing how to find the evidence of vulnerabilities in Verilog HDL.We first introduce a Unified Verilog Property Graph (VeriPG) that explicitly models parallel semantics by combining AST, CFG, and DDG.Furthermore, a semantic validation mechanism is designed to constrain and filter the LLM-generated TDPs.By executing these validated TDPs on VeriPG, our method produces stable and deterministic detection results.Experiments demonstrate that VerilogLAVD improves the F1 score by 133% compared to LLM-based methods.Furthermore, the framework successfully identifies real-world hardware vulnerabilities in open-source hardware design repositories.The code and datasets of this study are available at https://github.com/ Chip-Security-Lab/VerilogLAVD Xiang Long, Yingjie Xia, Li Kuang, Yao Wan 0001 |
ACL (1) | 2 |
| 2026 | LLMUpdater: Automatic comment synchronization via edit model guided LLMs
Haiyang Yang, Qi Xie 0010, Shixiang Cai, Li Kuang, Yingjie Xia |
Empir. Softw. Eng. | 6 |
| 2026 | LETA: A Lattice-Based Efficient and Traceable Privacy-Preserving Batch Authentication Scheme for Vehicle Platoon in VANETsabstractVehicle platoon (VP), as a typical form of traffic cooperation, can significantly enhance traffic efficiency and safety in Vehicular Ad hoc Networks (VANETs). However, malicious vehicles in VP poses a severe threat to the security of entire VP, requiring to be efficiently traced by identity authentication. In this paper, we propose a lattice-based efficient and traceable privacy-preserving batch authentication scheme for vehicle platoon in VANETs, named LETA. First, we design a dynamic VP identity structure VPD-Tree which is constructed based on hash tree and pseudonyms of vehicles to preserve privacy. Then, an aggregate signature is constructed based on VPD-tree and modular lattice for secure and efficient batch authentication of VP. Finally, Zero-Knowledge Proofs (ZKP) is applied on the VPD-Tree structure to anonymously and efficiently trace the malicious vehicles of VP. Security analysis shows that LETA achieves stronger security guarantees, thereby offering a more secure solution than existing approaches. Moreover, performance evaluations show that LETA achieves lower computation and communication overheads through the VPD-tree structure and efficient batch authentication scheme. Yingjie Xia, Xuejiao Liu 0002, Zhiquan Liu 0001, Zhen Guo 0003, Zhe Liu 0001, Liming Fang 0001 |
IEEE J. Sel. Areas Commun. | 2 |
| 2026 | ECroA: Efficient Cross-Domain Authentication in Dynamic Digital Twins of Wireless Industrial IoTabstractThe development of wireless industrial internet of things (IIoT) has facilitated the widespread deployment of dynamic digital twins across multiple factories and regions. However, the wireless communication between digital twins and their physical devices takes place over public channels, and the dynamic behavior of digital twins makes them susceptible to forgery and impersonation. As a result, data exchanged between digital twins is exposed to risks of tampering, fabrication and corruption. Moreover, real-time data sharing among a large number of digital twins incurs significant computational and communication overheads, thereby necessitating a secure and efficient cross-domain authentication scheme to ensure trustworthy identities and data integrity. To address these challenges, we propose ECroA, an efficient cross-domain authentication in dynamic digital twins of wireless IIoT. Specifically, we employ a distributed key management mechanism to establish trust across domains while reducing communication overhead. Additionally, we design an identity-based signature and pseudonym-based authentication scheme to provide conditional privacy preservation, ensuring secure data exchange without exposing sensitive data. To further enhance efficiency, we integrate signature aggregation and batch verification techniques, significantly reducing computational costs for large-scale cross-domain authentication. Performance analysis demonstrates that ECroA effectively enhances authentication efficiency while ensuring secure cross-domain authentication in dynamic digital twins of wireless IIoT. Yingjie Xia, Xuejiao Liu 0002, Yidan Yin |
IEEE J. Sel. Areas Commun. | 1 |
| 2026 | MCD4SR: Multimodal collaborative denoising with modality balancing for sequential recommendation
Xin Zhang 0079, Yinzhuo Chen, Shengan Wang, Dongjing Wang, Yingjie Xia, Sijie Niu, Butian Huang, Yuyu Yin |
Knowl. Based Syst. | 6 |
| 2026 | A Secure Self-Sovereign Identity Scheme Based on Soulbound Token and Trusted Personal Data SpaceabstractSelf-Sovereign Identity (SSI) aims to enable a physical-world persona to autonomously own and manage his digital identity, which serves as a mapping of his real-world identity into cyberspace and has become an essential paradigm in trusted digital identity management. Although current state-of-the-art SSI schemes provide security properties such as Sybil-resistance, privacy-preserving, or regulation, they typically employ static identifier(s) to represent a persona's identity, rendering them vulnerable to identity forgery or theft, and they still fail to simultaneously address the problems of privacy leakage, Sybil attacks, and inadequate regulation. To overcome these limitations, we propose a secure self-sovereign identity scheme that integrates SoulBound Token (SBT) and a trusted personal data space. In our design, each physical-world persona has a unique, individual-specific identity engine generated from verified biometric data and metadata, with a non-transferable SBT bound to it via a smart contract, thereby enhancing resistance to Sybil attacks and making the identity misappropriation significantly harder. To further prevent an identity from being linked by service providers, we construct a distinct child identity engine for each service using behavioral history stored in a trusted personal data space. Moreover, we design a conditional and collaborative regulation mechanism that ensures malicious digital identities and service providers are traceable and subject to supervision. Finally, we perform a security and performance analysis to demonstrate that our scheme is practical, secure, and offers several advantages over state-of-the-art solutions. Qiuyun Lyu, Xiaoqi Lou, Butian Huang, Yingjie Xia |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2026 | ESGA: Efficient and Secure Geo-Aware Certificate Revocation for Blockchain-Based Multi-Domain VANETsabstractIn vehicular ad-hoc networks (VANETs), certificate revocation is a critical technique to synchronize the invalid vehicle information accurately and timely across the network, especially in scenarios involving multiple edge domains. In order to achieve efficient and privacy-preserving certificate revocation list (CRL) synchronization, we propose ESGA, an efficient and secure geo-aware certificate revocation scheme for blockchain-based multi-domain VANETs. Since invalid vehicles are driving within specific edge domains during certain time periods, we design a hierarchical Geo-CRL structure comprising a blockchain-based global CRL and geographically distributed local CRLs. Furthermore, we design an adaptive geo-aware synchronization mechanism which enables timely local CRLs update, while ensuring periodic or event-driven updates of the global CRL on the blockchain. Finally, to protect the CRL information from pseudonym linking attacks, we design linkage values chain-based certificate revocation method for privacy preservation. The experimental results demonstrate that ESGA can significantly improve CRL synchronization efficiency with lower CRL query latency and storage overhead, and to protect the privacy of certificate revocation in blockchain-based multi-domain VANETs. Yingjie Xia, Tiancong Cao, Xuejiao Liu 0002 |
IEEE Trans. Mob. Comput. | 1 |
| 2025 | CSTree-SRI: Introspection-Driven Cognitive Semantic Tree for Multi-Turn Question Answering over Extra-Long ContextsabstractZhaowen Wang, Xiang Wei, Kangshao Du, Yiting Zhang, Libo Qin, Yingjie Xia, Li Kuang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Kangshao Du, Libo Qin 0001, Yingjie Xia, Li Kuang |
ACL (1) | 6 |
| 2025 | Human Pose Estimation Method Based on Top-Down View Fisheye Images
Quanyuan Chen, Yingjie Xia, Qun Xie, Jinping Li |
ICIG (2) | 4 |
| 2025 | Time-Frequency Domain-Based No-Reference Algorithm for Image Blurriness Evaluation
Hao Ning, Yingjie Xia, Qun Xie, Jinping Li |
ICIG (2) | 3 |
| 2025 | Video Stabilization Based on MeshFlow Motion Model in Dynamic and Complex Scenes
Hao Ning, Yingjie Xia, Qun Xie, Jinping Li |
ICIG (1) | 4 |
| 2025 | EvoAPR: Enhancing Large Language Models for Automatic Program Repair with Genetic Algorithm and Dynamic LoRAabstractAutomated Program Repair (APR) aims to automate the patch generation for buggy code and is vital in software devel-opment and maintenance. While large language models (LLMs) excel in various tasks, our empirical study shows they still face challenges in APR. LLMs take infilling templates with different qualities as input, and low-quality templates may misguide LLMs in constantly generating incorrect patches. Additionally, LLMs lack project-specific knowledge and struggle to leverage bug context fully. Therefore, we propose EvoAPR, which integrates the genetic algorithm and dynamic LoRA technique to enhance LLMs for better APR. First, to generate high-quality infilling templates, buggy codes are encoded to genetic representations, and genetic operators and multi-angle evaluation are designed to produce better templates. Then, to effectively utilize bug context, a bug-context-aware LoRA fine-tuning method is proposed to fuse various bug contexts by dynamically activating certain blocks of the LoRA module according to the bug index. Finally, these high-quality infilling templates are fed the fine-tuned LLMs for patch generation. Experimental results demonstrate that EvoAPR significantly improves LLMs' performance, surpassing state-of-the-art methods. Qingyang Yan, Weihuan Min, Li Kuang, Yingjie Xia |
ICWS | 6 |
| 2025 | SecV: LLM-based Secure Verilog Generation with Clue-Guided Exploration on Hardware-CWE Knowledge GraphabstractVerilog is specified as the primary Register Transfer Level (RTL) hardware description language, which designs the logical functions between registers for digital circuit systems. Recently, there emerges much cutting-edge research in leveraging Large Language Models (LLMs) to generate Verilog, aiming at effectively reducing errors and costs in the logic design of chips. However, these works mainly focus on logical correctness or PPA (Power, Performance, Area) measurement of the generated results, while neglecting the security problems in Verilog. In this study, we propose SecV, a novel and unified framework to generate secure Verilog by clue-guided exploration on Common Weakness Enumeration (CWE) knowledge graph (KG) for chips. First, the builder of the KG utilizes the instance-adapted chain of thought (COT) to extract entities and their relationships from raw Hardware-CWE corpora. Then, a fine-tuned BERT model is employed to verify the Hardware-CWE KG and collaborate with builder iteratively to achieve the precise KG. Based on Hardware-CWE KG, a clue-guided graph exploration paradigm is designed to facilitate collaborative inference of knowledge to generate secure Verilog by LLMs. Experiments demonstrate that SecV achieves 82.6% secure Verilog code without specified CWE in the generated functionally correct Verilog, with superior performance of a 21.7% performance improvement compared to SOTA. Fanghao Fan, Yingjie Xia, Li Kuang |
IJCAI | 2 |
| 2025 | APIMig: A Project-Level Cross-Multi-Version API Migration Framework Based on Evolution Knowledge GraphabstractAPI migration is essential for software maintenance due to the rapid evolution of third-party libraries where API elements may change continuously through updates. There are two main challenges for API migration at the project level, especially across multiple versions: 1) lack of specific library evolution knowledge across multi-version; 2) difficulty in identifying the chain of changes at the project level. This paper proposes a project-level cross-multi-version API migration framework APIMig. We first construct an API evolution knowledge graph (KG) to capture changes between adjacent library versions and then derive coherent cross-version API evolution knowledge by KG reasoning. Second, we design a chain exploration algorithm to track the chain of changes and aggregate the affected code segments. Finally, a large language model is employed in completing API migration by providing the API evolution knowledge and the chain of changes. We construct an evolution KG for the Lucene library from version 4.0.0 to 10.1.0 and evaluate our approach through project migration pairs that depend on different major versions. Our framework shows improvements over the baseline in migrating projects across 7 major versions, achieving average increases of 16.52% in CodeBLEU scores and 28.49% in VCEU scores in GPT-4o. Li Kuang, Qi Xie 0010, Haiyang Yang, HaoYue Kang, Yingjie Xia |
IJCAI | 7 |
| 2025 | DLCoG: A Novel Framework for Dual-Level Code Comment Generation Based on Semantic Segmentation and In-Context LearningabstractIn large software projects with collaborative development, comprehensive code comments are crucial for code readability and maintainability. Code comments mainly include method comments and inline comments, where the former describes the functionality globally, and the latter describes the implementation details locally. Existing methods typically generate these two kinds of comments with specific locations independently, which results in weak correlations between comments and code context, as well as high model inference costs due to long token inputs. To address these issues, we define the combination of inline comments and method comments as Dual-Level Code Comments. We formulate the novel task of automatically generate dual-level code comments based on given code and propose an approach named DLCoG (DualLevel Code Comment Generation) to automate this task. First, a Semantic Segmentation and Identification multi-task model based on CodeBERT, termed Se2Iden (Semantic Segmentation and Identification model), is proposed to identify code segments requiring inline comments. Next, we retrieve similar samples to adopting the in-context learning paradigm, which can enhance the generation quality of large language models (LLMs) in specific domains. Finally, the LLM is guided to generate duallevel code comments using Chain-of-Thought (CoT) prompts that first produce inline comments, followed by method comments. We manually constructed a high-quality clean Java dataset consisting of*> based on open-source Java projects by (i) determining comments type and (ii) manually associating inline comments with their corresponding code. Then, we trained a multi-task learning model based on CodeBERT to automatically take the two steps needed, termed ICSA (Inline Comment Classification and Scope Association), thus to expand to a dataset containing 80k dual-level code comments. Experimental results on clean and extended datasets show that DLCoG outperforms all baselines by substantial margins. The contextual information provided by DLCoG can effectively improve the inline comments generated by LLM. Coordinated generation of dual-level comment also brings effective improvements to method comments, which is particularly significant when there are few contextual examples. Our work fills the long-standing gap in the dual-level code comment generation field, and can provide insights for future research in this direction. We provide open-source datasets and source code for future research. Haiyang Yang, Qingyang Yan, Weihuan Min, Zhao Wei, Li Kuang, Yingjie Xia |
ICPC | 8 |
| 2025 | FCLLM-DT: Enpowering Federated Continual Learning With Large Language Models for Digital-Twin-Based Industrial IoTabstractThe Industrial Internet of Things (IIoT) represents a sophisticated technology designed to enhance production management and predict output in industrial settings, including machinery fault diagnostics. The precision of fault diagnosis is contingent upon the training efficacy of diagnostic models and their interoperability with models from other industrial facilities. Nonetheless, several critical challenges persist in maintaining these diagnostic models: 1) machinery sensors may generate abnormal data, resulting in suboptimal quality in model training; 2) sensor malfunctions may lead to interruptions in continuous data flow, thus impeding model training; and 3) collaborative interactions with other factories aiming at improving model performance may pose risks of privacy breaches. In this study, we introduce the FCLLM-DT scheme, which integrates the digital twin (DT) methodology to create a physical model of bearing for fixing abnormal sensor data. Additionally, retrieval-augmented generation (RAG)-assisted large language models (LLMs) are utilized to generate virtual datasets in instances of sensor failure. Moreover, for IIoT applications across distributed industrial environments, federated continual learning (FCL) is employed to enhance global model training by aggregating localized models from diverse facilities, thereby improving the accuracy of bearing fault diagnosis while safeguarding data privacy. The experiments on the accuracy of DT for abnormal data fix, RAG-assisted LLM for virtual data generation, and FCL for bearing fault diagnosis are conducted in comparison with three alternative methods across two datasets. The results indicate that our proposed scheme surpasses existing methods in both the enhancement of sensing data quality and the accuracy of bearing fault diagnosis. Yingjie Xia, Yunxiao Zhao, Li Kuang, Xuejiao Liu 0002, Ji Hu 0002, Zhiquan Liu 0001 |
IEEE Internet Things J. | 1 |
| 2025 | Personalized Privacy Preserving for Spatial Crowdsourcing by Reinforcement Learning in VANETsabstractSpatial crowdsourcing is widely used in various applications in vehicular ad hoc networks (VANETs), such as navigation system, traffic control, and event reporting. However, exposing and disclosing the spatial-temporal features of vehicles in the crowdsourcing services will definitely raise serious privacy issues. Existing unified privacy-preserving strategy for vehicles will cause excessive or insufficient preservation, thus result in relatively low-quality services. To solve these problems, we propose personalized privacy preserving for spatial crowdsourcing by reinforcement learning in VANETs. First, we propose a multifactor-based personalized privacy-preserving model to adjust the vehicles’ privacy-preserving level in the scenario of spatial crowdsourcing. And, we employ reinforcement learning to dynamically adjust the model in different situations. Furthermore, we propose an optimal local differential privacy mechanism to maintain the optimal tradeoff between data privacy and data utility, which can achieve personalized privacy preserving in the task allocation. We conduct extensive simulations with the real-world traffic trajectory dataset T-drive, and use the Q-learning algorithm to dynamically adjust the model. The experiments demonstrate that our scheme can enhance data utility by 78.5% with personalized privacy settings. Yingjie Xia, Tiancong Cao, Xuejiao Liu 0002, Zhiquan Liu 0001 |
IEEE Internet Things J. | 1 |
| 2025 | PPDR: A Privacy-Preserving Dual Reputation Management Scheme in Vehicle PlatoonabstractVehicle platoon has attracted much attention in recent years for its benefits in improving traffic efficiency, road safety, and energy consumption. In a vehicle platoon, vehicles can take on the roles of either a leading vehicle or a following vehicle, depending on the need. Selecting a reliable leading vehicle is crucial to improve the reliability of the vehicle platoon, and reputation management plays a vital role in this selection process. However, the existing reputation management schemes do not distinguish the leading vehicle's reputation value (LV-Reputation value) and the following vehicle's reputation value (FV-Reputation value), and merge them into a single reputation value, which exposes the reputation management scheme to the single reputation attack, an attack identified for the first time in this work. Additionally, some schemes fail to preserve the vehicle reputation value privacy, feedback score privacy, or identity privacy, and some overlook the security of the reputation management scheme. To address these issues, we propose a Privacy-Preserving Dual Reputation (PPDR) management scheme in vehicle platoon. The PPDR scheme separates a vehicle's reputation value into LV-Reputation value and FV-Reputation value, effectively mitigating the single reputation attack. Furthermore, under the premise of preserving reputation value privacy, feedback score privacy, and identity privacy, the PPDR scheme employs a score difference calculation algorithm to support the weighted average mechanism and the dual reputation management mechanism, which can significantly improve the accuracy of reputation management scheme. It also provides strong security with acceptable computation communication overheads. Comprehensive theoretical analysis and simulation evaluation have been carried out, and the results demonstrate that the PPDR scheme is significantly superior than the existing schemes in several aspects. Zhiquan Liu 0001, Yingjie Xia, Zhen Guo 0003, Gao Liu, Leping Li, Jianfeng Ma 0001 |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2025 | DRL-APG: Deep Reinforcement Learning Based Adaptive Policy Generation for Accurate and Secure Data Sharing in VANETsabstractWith the rapid expansion of vehicular ad-hoc networks (VANETs), the dynamic traffic environment raises concerns regarding accurate and secure data sharing. Unauthorized entities may exploit vulnerabilities to access sensitive information within shared data. To address these challenges, we propose DRL-APG, a deep reinforcement learning based adaptive policy generation scheme, to enable smarter security policies that can better handle complex changes in data sharing among vehicles. DRL-APG adopts hybrid modeling to capture environment states and fine-grained policy optimization using an ARIMA (Autoregressive Integrated Moving Average)-based reward mechanism to generate adaptive policies. Extensive simulations demonstrate DRL-APG outperforms existing schemes in different kinds of VANET situations including peak/off-peak hours and varying speed limits. In traffic congestion scenario across three simulation typical road networks, average accident zone speed increases 21.7%, 25.6% and 27.8% respectively after data sharing with the policy generated by our proposed DRL-APG, compared to existing schemes. Tiancong Cao, Xuejiao Liu 0002, Yingjie Xia |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | CF-SOLT: Real-time and accurate traffic accident detection using correlation filter-based tracking
Yingjie Xia, Nan Qian, Zheming Cai |
Image Vis. Comput. | 1 |
| 2024 | ARSL-V: A risk-aware relay selection scheme using reinforcement learning in VANETs
Xuejiao Liu 0002, Chuanhua Wang, Yingjie Xia |
Peer Peer Netw. Appl. | 4 |
| 2024 | Learning Appearance-Motion Synergy via Memory-Guided Event Prediction for Video Anomaly DetectionabstractClassic unsupervised anomaly detection learns normative patterns from normal behavior and assumes that unforeseen anomalous behavior will result in significant prediction deviations. However, anomaly detection in specific situations faces challenges in detecting ambiguous behavior in which the abnormal representation is not particularly intuitive. Existing anomaly detection approaches perform poorly for ambiguous behavior due to limited normative representational capacity, resulting in a narrow normality gap. We observe that the ambiguity of behavior comes from the contradiction between the properties of appearance and motion modalities. In this paper, we propose a novel memory-guided autoencoder named appearance-motion synergy autoencoder to detect anomalous behavior by event prediction. To address the above challenge, we leverage the synergy of the normative appearance-motion modalities to strengthen the representation of normative patterns and improve the detection of ambiguous behavior. Specifically, we design the memory networks with dynamic fusion mechanisms to integrate the correlated appearance-motion information and to remember normal patterns. A consistency measurement unit is designed to optimize the consistency of normative appearance-motion features via a joint distribution measurement pool. A larger normality gap in detecting ambiguous behavior in our approach enhances the abnormal detection capability. Extensive experiments demonstrate our superiority in detecting anomalous behavior. Chongye Guo, Yingjie Xia, Guorui Feng |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | EPP-GAS: An Efficient and Privacy-Preserving Cross Trust-Domain Group Authentication Scheme for Vehicle Platoon Based on BlockchainabstractThe vehicle platooning (VP) going across different regions together, plays an important role in improving traffic efficiency. Then how to perform authentication efficiently for the vehicle platoons becomes the basic requirements in the Internet of Vehicles, especially cross different trust domains. In this paper, we propose an efficient and privacy-preserving cross trust-domain group authentication scheme for VP based on blockchain, named EPP-GAS. We design an extended blockchain model, which encodes the VP authentication parameters into the BM-Tree block, to support the sharing of VP authentication parameters cross multiple trust-domains. Then, we construct group pseudonym based on Shamir’s threshold scheme, to realize anonymous group authentication efficiently. Also, we propose a group authentication scheme with batch identity authentication through dynamic adaption of VP adjustment and efficient re-authentication in the case of authentication failure. Compared with the existing authentication schemes, our scheme has better performance in terms of security, authentication efficiency, communication overhead, and increase efficiency for cross trust-domain group authentication by 31% on average and reduce the communication overhead by 43.9% on average. Yingjie Xia, Xuejiao Liu 0002, Qiang Zhong |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2024 | Retargeting HR Aerial Photos Under Contaminated Labels With Application in Smart NavigationabstractRetargeting aims to shrink a photo wherein the perceptually prominent regions are appropriately kept. In practice, optimally shrinking a high resolution (HR) aerial photo is a useful tool for smart navigation. Nowadays, vehicle drivers’ path planning is generally guided by an HR aerial photo recommended by a navigation App like Google Maps. Owing to the limited and various resolution of vehicle displays, we have to retarget each original HR aerial photo accordingly, wherein the navigation-aware regions can be well preserved. In practice, HR aerial photo retargeting is non-trivial due to three challenges: 1) the rich number of internal objects and their complex spatial layouts, 2) deriving the region-level semantics from potentially contaminated image labels, and 3) the inefficiency of retargeting each HR aerial photo with millions of pixels. To handle these problems, we propose a novel HR aerial photo retargeting pipeline that can intelligently avoid the negative effects from incorrect image labels. The key is a noise-tolerant hashing algorithm that converts image-level semantics into the hash codes corresponding to different regions, which guides the HR aerial photo shrinking. More specifically, for each HR aerial photo, we extract visually/semantically salient object patches inside it. To explicitly encode their spatial layout, we construct a graphlet by linking the spatially adjacent object patches into a small graph. Subsequently, a binary matrix factorization (MF) is designed to exploit the underlying semantics of these graphlets, wherein three attributes: i) binary hash codes learning, ii) noisy labels refinement, iii) deep image-level semantics, are collaboratively encoded. Such binary MF can be solved iteratively and each graphlet is subsequently converted into the binary hash codes. Finally, the hash codes corresponding to graphlets within each HR aerial photo are utilized to learn a Gaussian mixture model (GMM) that optimizes the HR aerial photo retargeting. During the experimental validation, we compiled a smart navigation dataset including 132743 planned paths annotated from 10132 HR aerial photos, based on which comparative study has demonstrated the superiority of our method. Bing Tu, Yingjie Xia |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2024 | CD-BASA: An Efficient Cross-Domain Batch Authentication Scheme Based on Blockchain With Accumulator for VANETsabstractAuthentication plays an important role in verifying the integrity of messages and the identity of message senders in vehicular ad-hoc networks (VANET) applications. A significant challenge in expansive regions pertains to the efficient authentication of multiple vehicles traversing trust domains. This paper introduces a novel cross-domain batch authentication scheme for VANETs, denominated as CD-BASA, leveraging blockchain technology in conjunction with accumulator. The CD-BASA scheme is built upon blockchain with accumulator and interplanetary file systems for streamlined storage and batch computation. Subsequently, a proficient cross-domain batch authentication protocol for CD-BASA is devised, encompassing the establishment of domain-specific and global parameters, which greatly reduces the identity verification cost for multiple vehicles. To further minimize the overhead associated with batch validity proofing, a novel multi-accumulator batch proof algorithm is designed to greatly reduce the proof cost of multiple vehicle identities. Experimental evaluations and theoretical security analysis show that CD-BASA has better performance than other existing schemes, in terms of cross-domain batch authentication overhead (improving at least 18.2%) and accumulator proof overhead (reducing at least 69.9%). Qiang Zhong, Yingjie Xia, Xuejiao Liu 0002 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2023 | RLID-V: Reinforcement Learning-Based Information Dissemination Policy Generation in VANETsabstractCiphertext policy attribute-based encryption (CP-ABE) is popularly used to implement secure and accurate access control of disseminated information in vehicular ad hoc networks (VANETs). Nevertheless, how to improve the policy generation of CP-ABE for accurate information dissemination in the dynamic VANETs remains a challenge, as there are several access control policies rising from moving vehicles and road side units (RSUs) with different sensing boarder regarding to a specific event, such as moving vehicles and road side units (RSUs). To solve this problem, this paper proposes a reinforcement learning-based information dissemination policy generation scheme in VANETs, named RLID-V. The scheme firstly combines multiple attribute-based access control policies and resolves policy conflicts between vehicles and RSUs. Then, a manual feedback policy construction method is designed by applying decision tree to the collected feedback from all receivers. Finally, we employ reinforcement learning to dynamically update the confidence weights of different policy sources. The experiments are conducted in two classic VANETs scenarios, traffic guidance and accident warning, demonstrating that RLID-V achieves better performance in the accuracy and effectiveness of information dissemination compared with three existing schemes. Otherwise, RLID-V outperforms the compared schemes in robustness with 20% error feedback and takes a negligible cost of less than 1% of the overall delay overhead for policy generation. Yingjie Xia, Xuejiao Liu 0002, Jing Ou, Oubo Ma |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2022 | SCMP-V: A secure multiple relays cooperative downloading scheme with privacy preservation in VANETs
Xuejiao Liu 0002, Chuanhua Wang, Wei Chen 0147, Yingjie Xia, Gaoxiang Zhu |
Peer-to-Peer Netw. Appl. | 4 |
| 2022 | Security-Aware Information Dissemination With Fine-Grained Access Control in Cooperative Multi-RSU of VANETsabstractSecuring information dissemination is extremely important for various vehicular ad hoc networks (VANETs) applications. However, most of the applications require disseminating the critical information only to the authorized vehicles through V2I (Vehicle to Infrastructure) communications. Moreover, V2I communications usually suffer from incomplete information to vehicles in one RSU’s transmission range. Therefore, it is an even challenging task that how to ensure reliable dissemination of encrypted data to the recipient vehicles in multi-RSU settings. We propose a security-aware information dissemination scheme with fine-grained access control in cooperative multi-RSU of VANETs. Our proposed scheme uses ciphertext-policy attribute-based encryption (CP-ABE) to ensure confidential communication in a broadcasting way, which ensures that only the vehicles that satisfy the access control policy can have the ability to access the information; and we employ proxy re-encryption in the communication protocol to make sure the moving vehicles in high-speed can get the whole encrypted information. Performance analysis shows that our scheme can enable fine-grained access control for the broadcasted information, and make sure the vehicles receive reliable information in the whole disseminating process. And our scheme is applicable and efficient in various scenarios of information dissemination, especially in cooperative multi-RSU of VANETs. Xuejiao Liu 0002, Wei Chen 0147, Yingjie Xia |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | TRAMS: A Secure Vehicular Crowdsensing Scheme Based on Multi-Authority Attribute-Based SignatureabstractRecently, vehicular crowdsensing networks have attracted much attention because of their ability to provide efficient and convenient information services for the Internet of Vehicles. How to achieve on-demand message authentication and provide privacy protection of sensing vehicles are challenging in accurate sensing tasks. We propose a secure vehicular crowdsensing scheme based on multi-authority attribute-based signature (TRAMS), which allows the publisher to flexibly customize a fine-grained policy that the potential participants must satisfy and uses attribute-based signature to authenticate sensed messages while protecting the privacy of the sensing vehicle. Also, we propose a multi-authority key management scheme, which can improve vehicle-based sensing efficiency in the Internet of Vehicles. Performance analysis shows that our scheme can not only achieve massage authentication while protecting the privacy of the sensing vehicle, but also ensure fine-grained message authentication to meet the expectation of the publisher on demand. And compared with the single-authority schemes in vehicular communication, our multi-authority TRAMS can achieve efficient message authentication for vehicular crowdsensing applications which require timely task feedback. Xuejiao Liu 0002, Wei Chen 0147, Yingjie Xia, Renhao Shen |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | HDRS: A Hybrid Reputation System With Dynamic Update Interval for Detecting Malicious Vehicles in VANETsabstractThe reputation-based scheme is a promising solution to prevent malicious behaviors in Vehicular Ad-hoc Networks (VANETs). However, traditional centralized reputation schemes are not suited for distributed networks, while decentralized reputation schemes are vulnerable to malicious vehicles spreading false messages. Most of these schemes assume that the behavior of vehicles can be accurately measured as reputation from the communication, ignoring that malicious vehicles may behave intelligently to avoid being detected. In this paper, we propose a hybrid reputation system (HDRS) which allows vehicles and roadside units (RSU) to complete reputation evaluations separately and provide references to each other. HDRS utilizes a reliability evaluation module to filter out unreliable calculation results and reference records. Furthermore, HDRS includes a dynamic adjustment mechanism for the reputation update interval, employing Analytic Hierarchy Process (AHP) and reliability evaluation results to resist intelligent attacks. Simulation results illustrate that HDRS can maintain a high detection rate and low false-positive rate for detecting malicious vehicles in different environments. Compared with existing schemes, HDRS increases the detection rates of collusion and intelligent attacks by 30% and 16%, respectively. Xuejiao Liu 0002, Oubo Ma, Wei Chen 0147, Yingjie Xia |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2022 | STAP: A Spatio-Temporal Correlative Estimating Model for Improving Quality of Traffic DataabstractWith the rapid development of intelligent transportation systems (ITS), traffic data plays a more and more important role. Low quality traffic data has become a challenging issue in the implementation of ITS. Inspired by the fact that traffic data have strong spatio-temporal correlation, we propose a quality improving model for traffic data, which correlates spatial and temporal features to fix abnormal data. We call it STAP, a spatio-temporal correlative estimating model which firstly proposes an anomalies detection algorithm based on an improved Random Forest model, and then classifies traditional features and extracts spatial and temporal features respectively. Finally the model proposes an XGboost-based data estimation algorithm to fix abnormal data. We conduct experiments on real traffic data collected from a big China city, Changsha, and the results show that the STAP model is effective in improving data quality. Yingjie Xia, Jing Ou |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2021 | SE-VFC: Secure and Efficient Outsourcing Computing in Vehicular Fog ComputingabstractFog-aided computing is an emerging computing paradigm developed for providing various computation services to end users in vehicular fog computing. However, fog vehicles may be untrusted, and results from malicious operations may cause serious accidents. Therefore, we propose a secure and efficient outsourcing computing scheme in vehicular fog computing (SE-VFC), which performs outsourcing computing through fog vehicles with computing resources, and combines lightweight Boneh-Lynn-Shacham (BLS) signature and group signature to achieve batch anonymous authentication of fog vehicles while protecting their privacy in multiple outsourcing tasks. We also verify the correctness of the outsourcing computing results. Security analysis shows that our scheme can authenticate the fog vehicles in the outsourcing computation and preserve their privacy, it can also trace the malicious ones if necessary. Compared with the existing schemes, our scheme has relatively low communication and computation overhead, thus it is quite efficient for batch authentication especially in multiple computing tasks. Extensive simulation results verify the effectiveness and practicality of our proposed scheme in vehicular fog computing. Xuejiao Liu 0002, Wei Chen 0147, Yingjie Xia, Chenghan Yang |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2018 | Secure and efficient querying over personal health records in cloud computing
Xuejiao Liu 0002, Yingjie Xia, Fengli Yang |
Neurocomputing | 2 |
| 2018 | Exploring Web Images to Enhance Skin Disease Analysis Under A Computer Vision FrameworkabstractTo benefit the skin care, this paper aims to design an automatic and effective visual analysis framework, with the expectation of recognizing the skin disease from a given image conveying the disease affected surface. This task is nontrivial, since it is hard to collect sufficient well-labeled samples. To address such problem, we present a novel transfer learning model, which is able to incorporate external knowledge obtained from the rich and relevant Web images contributed by grassroots. In particular, we first construct a target domain by crawling a small set of images from vertical and professional dermatological websites. We then construct a source domain by collecting a large set of skin disease related images from commercial search engines. To reinforce the learning performance in the target domain, we initially build a learning model in the target domain, and then seamlessly leverage the training samples in the source domain to enhance this learning model. The distribution gap between these two domains are bridged by a linear combination of Gaussian kernels. Instead of training models with low-level features, we resort to deep models to learn the succinct, invariant, and high-level image representations. Different from previous efforts that focus on a few types of skin diseases with a small and confidential set of images generated from hospitals, this paper targets at thousands of commonly seen skin diseases with publicly accessible Web images. Hence the proposed model is easily repeatable by other researchers and extendable to other disease types. Extensive experiments on a real-world dataset have demonstrated the superiority of our proposed method over the state-of-the-art competitors. Yingjie Xia, Lei Meng 0001, Yan Yan 0002, Liqiang Nie, Xuelong Li 0001 |
IEEE Trans. Cybern. | 1 |
| 2018 | Toward Personalized Activity Level Prediction in Community Question Answering WebsitesabstractCommunity Question Answering (CQA) websites have become valuable knowledge repositories. Millions of internet users resort to CQA websites to seek answers to their encountered questions. CQA websites provide information far beyond a search on a site such as Google due to (1) the plethora of high-quality answers, and (2) the capabilities to post new questions toward the communities of domain experts. While most research efforts have been made to identify experts or to preliminarily detect potential experts of CQA websites, there has been a remarkable shift toward investigating how to keep the engagement of experts. Experts are usually the major contributors of high-quality answers and questions of CQA websites. Consequently, keeping the expert communities active is vital to improving the lifespan of these websites. In this article, we present an algorithm termed PALP to predict the activity level of expert users of CQA websites. To the best of our knowledge, PALP is the first approach to address a personalized activity level prediction model for CQA websites. Furthermore, it takes into consideration user behavior change over time and focuses specifically on expert users. Extensive experiments on the Stack Overflow website demonstrate the competitiveness of PALP over existing methods. Zhenguang Liu, Yingjie Xia, Qi Liu 0049, Qinming He, Chao Zhang 0014, Roger Zimmermann |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2017 | FastShrinkage: Perceptually-aware Retargeting Toward Mobile PlatformsabstractRetargeting aims at adapting an original high-resolution photo/video to a low-resolution screen with an arbitrary aspect ratio. Conventional approaches are generally based on desktop PCs, since the computation might be intolerable for mobile platforms (especially when retargeting videos). Besides, only low-level visual features are exploited typically, whereas human visual perception is not well encoded. In this paper, we propose a novel retargeting framework which fast shrinks photo/video by leveraging human gaze behavior. Specifically, we first derive a geometry-preserved graph ranking algorithm, which efficiently selects a few salient object patches to mimic human gaze shifting path (GSP) when viewing each scenery. Afterward, an aggregation-based CNN is developed to hierarchically learn the deep representation for each GSP. Based on this, a probabilistic model is developed to learn the priors of the training photos which are marked as aesthetically-pleasing by professional photographers. We utilize the learned priors to efficiently shrink the corresponding GSP of a retargeted photo/video to be maximally similar to those from the training photos. Extensive experiments have demonstrated that: 1) our method consumes less than 35ms to retarget a 1024 × 768 photo (or a 1280 × 720 video frame) on popular iOS/Android devices, which is orders of magnitude faster than the conventional retargeting algorithms; 2) the retargeted photos/videos produced by our method outperform its competitors significantly based on the paired-comparison-based user study; and 3) the learned GSPs are highly indicative of human visual attention according to the human eye tracking experiments. Zhenguang Liu, Zepeng Wang 0003, Rajiv Ratn Shah, Yingjie Xia, Yi Yang 0001, Xuelong Li 0001 |
ACM Multimedia | 5 |
| 2017 | Learning for classification of traffic-related object on RGB-D data
Yingjie Xia, Xingmin Shi |
Multim. Syst. | 1 |
| 2017 | Perceptually Guided Photo RetargetingabstractWe propose perceptually guided photo retargeting, which shrinks a photo by simulating a human's process of sequentially perceiving visually/semantically important regions in a photo. In particular, we first project the local features (graphlets in this paper) onto a semantic space, wherein visual cues such as global spatial layout and rough geometric context are exploited. Thereafter, a sparsity-constrained learning algorithm is derived to select semantically representative graphlets of a photo, and the selecting process can be interpreted by a path which simulates how a human actively perceives semantics in a photo. Furthermore, we learn the prior distribution of such active graphlet paths (AGPs) from training photos that are marked as esthetically pleasing by multiple users. The learned priors enforce the corresponding AGP of a retargeted photo to be maximally similar to those from the training photos. On top of the retargeting model, we further design an online learning scheme to incrementally update the model with new photos that are esthetically pleasing. The online update module makes the algorithm less dependent on the number and contents of the initial training data. Experimental results show that: 1) the proposed AGP is over 90% consistent with human gaze shifting path, as verified by the eye-tracking data, and 2) the retargeting algorithm outperforms its competitors significantly, as AGP is more indicative of photo esthetics than conventional saliency maps. Yingjie Xia, Richang Hong, Liqiang Nie, Yan Yan 0002, Ling Shao 0001 |
IEEE Trans. Cybern. | 1 |
| 2017 | Weakly Supervised Multimodal Kernel for Categorizing Aerial PhotographsabstractAccurately distinguishing aerial photographs from different categories is a promising technique in computer vision. It can facilitate a series of applications, such as video surveillance and vehicle navigation. In this paper, a new image kernel is proposed for effectively recognizing aerial photographs. The key is to encode high-level semantic cues into local image patches in a weakly supervised way, and integrate multimodal visual features using a newly developed hashing algorithm. The flowchart can be elaborated as follows. Given an aerial photo, we first extract a number of graphlets to describe its topological structure. For each graphlet, we utilize color and texture to capture its appearance, and a weakly supervised algorithm to capture its semantics. Thereafter, aerial photo categorization can be naturally formulated as graphlet-to-graphlet matching. As the number of graphlets from each aerial photo is huge, to accelerate matching, we present a hashing algorithm to seamlessly fuze the multiple visual features into binary codes. Finally, an image kernel is calculated by fast matching the binary codes corresponding to each graphlet. And a multi-class SVM is learned for aerial photo categorization. We demonstrate the advantage of our proposed model by comparing it with state-of-the-art image descriptors. Moreover, an in-depth study of the descriptiveness of the hash-based graphlet is presented. Yingjie Xia, Zhenguang Liu, Liqiang Nie, Xuelong Li 0001 |
IEEE Trans. Image Process. | 1 |
| 2017 | Adaptive Multimedia Data Forwarding for Privacy Preservation in Vehicular Ad-Hoc NetworksabstractVehicular ad-hoc networks (VANETs) have drawn much attention of researchers. The vehicles in VANETs frequently join and leave the networks, and therefore restructure the network dynamically and automatically. Forwarded messages in vehicular ad-hoc networks are primarily multimedia data, including structured data, plain text, sound, and video, which require access control with efficient privacy preservation. Ciphertext-policy attribute-based encryption (CP-ABE) is adopted to meet the requirements. However, solutions based on traditional CP-ABE suffer from challenges of the limited computational resources on-board units equipped in the vehicles, especially for the complex policies of encryption and decryption. In this paper, we propose a CP-ABE delegation scheme, which allows road side units (RSUs) to perform most of the computation, for the purpose of improving the decryption efficiency of the vehicles. By using decision tree to jointly optimize multiple factors, such as the distance from RSU, the communication and computational cost, the CP-ABE delegation scheme is adaptively activated based on the estimation of various vehicles decryption overhead. Experimental results thoroughly demonstrate that our scheme is effective and efficient for multimedia data forwarding in vehicular ad-hoc networks with privacy preservation. Yingjie Xia, Wenzhi Chen, Xuejiao Liu 0002, Xuelong Li 0001, Yang Xiang 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2017 | Media Quality Assessment by Perceptual Gaze-Shift Patterns DiscoveryabstractQuality assessment is an indispensable technique in a large body of media applications, i.e., photo retargeting, scenery rendering, and video summarization. In this paper, a fully automatic framework is proposed to mimic how humans subjectively perceive media quality. The key is a locality-preserved sparse encoding algorithm that accurately discovers human gaze shifting paths from each image or video clip. In particular, we first extract local image descriptors from each image/video, and subsequently project them into the so-called perceptual space. Then, a nonnegative matrix factorization (NMF) algorithm is proposed that represents each graphlet by a linear and sparse combination of the basis ones. Since each graphlet is visually/semantically similar to its neighbors, a locality-preserved constraint is encoded into the NMF algorithm. Mathematically, the saliency of each graphlet is quantified by the norm of its sparse codes. Afterward, we sequentially link them into a path to simulate human gaze allocation. Finally, a probabilistic quality model is learned based on such paths extracted from a collection of photos/videos, which are marked as high quality ones via multiple Flickr users. Comprehensive experiments have demonstrated that: 1) our quality model outperforms many of its competitors significantly, and 2) the learned paths are on average 89.5% consistent with real human gaze shifting paths. Yingjie Xia, Zhenguang Liu, Yan Yan 0002, Roger Zimmermann |
IEEE Trans. Multim. | 1 |
| 2016 | Learning for Traffic State Estimation on Large Scale of Incomplete DataabstractTraffic state estimation is one of the very basic and most important tasks in intelligent transportation systems. Data collected from multiple types of traffic sensors are valuable sources to estimate traffic state. However, those data are always incomplete for the cases such as sensors broken, data missing, etc. The incomplete data can also be used to accurately estimate traffic state by data prediction and data fusion. Firstly, the missing data are predicted based on the data correlations extracted on large amount of historical taxi GPS data. Then, data collected from multiple sensors are fused to get more accurate results. Finally, computing of data fusion is parallelized. The main research focus is placed on learning for extracting the inherent spatio-temporal correlations of traffic data and fusing multimedia traffic data. Accuracy and efficiency experiments are conducted and the experimental results demonstrate that our method can process incomplete data and achieve accurate and real-time traffic state estimation. Yiyang Yao, Yingjie Xia, Zhenyu Shan, Zhengguang Liu |
ICMR | 2 |
| 2016 | Utilizing Sensor-Social Cues to Localize Objects-of-Interest in Outdoor UGVs
Yingjie Xia, Liqiang Nie, Wenjing Geng |
MMM (1) | 1 |
| 2016 | Big traffic data processing framework for intelligent monitoring and recording systems
Yingjie Xia, Xindai Lu |
Neurocomputing | 1 |
| 2016 | Special issue on big data driven Intelligent Transportation Systems
Yingjie Xia, Yuncai Liu |
Neurocomputing | 1 |
| 2016 | Multi-source alert data understanding for security semantic discovery based on rough set theory
Yiyang Yao, Chun Gan, Qian Kang, Xuejiao Liu 0002, Yingjie Xia |
Neurocomputing | 6 |
| 2016 | Toward solving the Steiner travelling salesman problem on urban road maps using the branch decomposition of graphs
Yingjie Xia, Mingzhe Zhu, Qian-Ping Gu, Xuelong Li 0001 |
Inf. Sci. | 1 |
| 2016 | SEMD: Secure and efficient message dissemination with policy enforcement in VANET
Xuejiao Liu 0002, Yingjie Xia, Wenzhi Chen, Yang Xiang 0001, Mohammad Mehedi Hassan, Abdulhameed Alelaiwi |
J. Comput. Syst. Sci. | 2 |
| 2016 | APPOS: An adaptive partial occlusion segmentation method for multiple vehicles tracking
Yingjie Xia, Xingmin Shi, Yuncai Liu |
J. Vis. Commun. Image Represent. | 2 |
| 2016 | Geometric discriminative features for aerial image retrieval in social media
Yingjie Xia, Ying Zhang 0047 |
Multim. Syst. | 1 |
| 2016 | Detection and classification of traffic lights for automated setup of road surveillance systems
Xingmin Shi, Yingjie Xia |
Multim. Tools Appl. | 3 |
| 2016 | Effectively identifying the influential spreaders in large-scale social networks
Yingjie Xia, Zhengchao Peng, Li She |
Multim. Tools Appl. | 1 |
| 2016 | Towards improving quality of video-based vehicle counting method for traffic flow estimation
Yingjie Xia, Xingmin Shi, Guanghua Song, Qiaolei Geng, Yuncai Liu |
Signal Process. | 1 |
| 2016 | A reduction of security notions in designated confirmer signatures
Yingjie Xia, Xuejiao Liu 0002, Fubiao Xia, Guilin Wang |
Theor. Comput. Sci. | 1 |
| 2016 | Perceptual Attributes Optimization for Multivideo SummarizationabstractNowadays, many consumer videos are captured by portable devices such as iPhone. Different from constrained videos that are produced by professionals, e.g., those for broadcast, summarizing multiple handheld videos from a same scenery is a challenging task. This is because: 1) these videos have dramatic semantic and style variances, making it difficult to extract the representative key frames; 2) the handheld videos are with different degrees of shakiness, but existing summarization techniques cannot alleviate this problem adaptively; and 3) it is difficult to develop a quality model that evaluates a video summary, due to the subjectiveness of video quality assessment. To solve these problems, we propose perceptual multiattribute optimization which jointly refines multiple perceptual attributes (i.e., video aesthetics, coherence, and stability) in a multivideo summarization process. In particular, a weakly supervised learning framework is designed to discover the semantically important regions in each frame. Then, a few key frames are selected based on their contributions to cover the multivideo semantics. Thereafter, a probabilistic model is proposed to dynamically fit the key frames into an aesthetically pleasing video summary, wherein its frames are stabilized adaptively. Experiments on consumer videos taken from sceneries throughout the world demonstrate the descriptiveness, aesthetics, coherence, and stability of the generated summary. Liqiang Nie, Richang Hong, Yingjie Xia, Dacheng Tao, Nicu Sebe |
IEEE Trans. Cybern. | 4 |
| 2016 | Weakly Supervised Multilabel Clustering and its Applications in Computer VisionabstractClustering is a useful statistical tool in computer vision and machine learning. It is generally accepted that introducing supervised information brings remarkable performance improvement to clustering. However, assigning accurate labels is expensive when the amount of training data is huge. Existing supervised clustering methods handle this problem by transferring the bag-level labels into the instance-level descriptors. However, the assumption that each bag has a single label limits the application scope seriously. In this paper, we propose weakly supervised multilabel clustering, which allows assigning multiple labels to a bag. Based on this, the instance-level descriptors can be clustered with the guidance of bag-level labels. The key technique is a weakly supervised random forest that infers the model parameters. Thereby, a deterministic annealing strategy is developed to optimize the nonconvex objective function. The proposed algorithm is efficient in both the training and the testing stages. We apply it to three popular computer vision tasks: 1) image clustering; 2) semantic image segmentation; and 3) multiple objects localization. Impressive performance on the state-of-the-art image data sets is achieved in our experiments. Yingjie Xia, Liqiang Nie, Yi Yang 0001, Richang Hong, Xuelong Li 0001 |
IEEE Trans. Cybern. | 1 |
| 2016 | Weakly Supervised Human Fixations PredictionabstractAutomatically predicting human eye fixations is a useful technique that can facilitate many multimedia applications, e.g., image retrieval, action recognition, and photo retargeting. Conventional approaches are frustrated by two drawbacks. First, psychophysical experiments show that an object-level interpretation of scenes influences eye movements significantly. Most of the existing saliency models rely on object detectors, and therefore, only a few prespecified categories can be discovered. Second, the relative displacement of objects influences their saliency remarkably, but current models cannot describe them explicitly. To solve these problems, this paper proposes weakly supervised fixations prediction, which leverages image labels to improve accuracy of human fixations prediction. The proposed model hierarchically discovers objects as well as their spatial configurations. Starting from the raw image pixels, we sample superpixels in an image, thereby seamless object descriptors termed object-level graphlets (oGLs) are generated by random walking on the superpixel mosaic. Then, a manifold embedding algorithm is proposed to encode image labels into oGLs, and the response map of each prespecified object is computed accordingly. On the basis of the object-level response map, we propose spatial-level graphlets (sGLs) to model the relative positions among objects. Afterward, eye tracking data is employed to integrate these sGLs for predicting human eye fixations. Thorough experiment results demonstrate the advantage of the proposed method over the state-of-the-art. Xuelong Li 0001, Liqiang Nie, Yi Yang 0001, Yingjie Xia |
IEEE Trans. Cybern. | 5 |
| 2015 | Weakly Supervised Random Forest for Multi-Label Image Clustering and SegmentationabstractClustering is a useful statistical tool in data mining and computer vision. Supervised information is introduced to improve the clustering performance. However, labeling each piece of data accurately is extremely expensive when the amount of data is huge. Existing supervised clustering methods handle the huge workload of labeling large amount of data by transferring the bag-level labels into the instance-level descriptors. However, each bag has only one label limits the application scope seriously. In this paper, we propose weakly supervised multi-label clustering, which allows to label a bag of data multiple labels. The key technique is a weakly supervised random forest which can calculate the model parameters with a deterministic annealing strategy to optimize the non-convex objective function. The proposed algorithm is applied to two typical applications, image clustering and segmentation problems. Impressive efficiency in both training and testing stages on the state-of-the-art image data sets is achieved in our experiments. Yingjie Xia, Wei Wei 0002 |
ICMR | 1 |
| 2015 | Biologically Inspired Media Quality ModelingabstractA successful quality model is indispensable in a rich variety of multimedia applications, e.g., image classification and video summarization. Conventional approaches have developed many features to assess media quality at both low-level and high-level. However, they cannot reflect the process of human visual cortex in media perception. It is generally accepted that an ideal quality model should be biologically plausible, i.e., capable of mimicking human gaze shifting as well as the complicated visual cognition. In this paper, we propose a biologically inspired quality model, focusing on interpreting how humans perceive visually and semantically important regions in an image (or a video clip). Particularly, we first extract local descriptors (graphlets in this work) from an image/frame. They are projected onto the perceptual space, which is built upon a set of low-level and high-level visual features. Then, an active learning algorithm is utilized to select graphlets that are both visually and semantically salient. The algorithm is based on the observation that each graphlet can be linearly reconstructed by its surrounding ones, and spatially nearer ones make a greater contribution. In this way, both the local and global geometric properties of an image/frame can be encoded in the selection process. These selected graphlets are linked into a so-called biological viewing path (BVP) to simulate human visual perception. Finally, the quality of an image or a video clip is predicted by a probabilistic model. Experiments shown that 1) the predicted BVPs are over 90% consistent with real human gaze shifting paths on average; and 2) our quality model outperforms many of its competitors remarkably. Meng Wang 0001, Liqiang Nie, Richang Hong, Yingjie Xia, Roger Zimmermann |
ACM Multimedia | 5 |
| 2015 | Formalizing computational intensity of big traffic data understanding and analysis for parallel computing
Yingjie Xia |
Neurocomputing | 1 |
| 2015 | Integrating 3D structure into traffic scene understanding with RGB-D data
Yingjie Xia, Weiwei Xu 0003, Xingmin Shi, Kuang Mao |
Neurocomputing | 1 |
| 2015 | Recognizing multi-view objects with occlusions using a deep architecture
Yingjie Xia, Weiwei Xu 0003, Zhenyu Shan, Yuncai Liu |
Inf. Sci. | 1 |
| 2015 | Vehicles overtaking detection using RGB-D data
Yingjie Xia, Xingmin Shi |
Signal Process. | 1 |
| 2015 | Scalable Constrained Spectral ClusteringabstractConstrained spectral clustering (CSC) algorithms have shown great promise in significantly improving clustering accuracy by encoding side information into spectral clustering algorithms. However, existing CSC algorithms are inefficient in handling moderate and large datasets. In this paper, we aim to develop a scalable and efficient CSC algorithm by integrating sparse coding based graph construction into a framework called constrained normalized cuts. To this end, we formulate a scalable constrained normalized-cuts problem and solve it based on a closed-form mathematical analysis. We demonstrate that this problem can be reduced to a generalized eigenvalue problem that can be solved very efficiently. We also describe a principled k-way CSC algorithm for handling moderate and large datasets. Experimental results over benchmark datasets demonstrate that the proposed algorithm is greatly cost-effective, in the sense that (1) with less side information, it can obtain significant improvements in accuracy compared to the unsupervised baseline; (2) with less computational time, it can achieve high clustering accuracies close to those of the state-of-the-art. Jianyuan Li, Yingjie Xia, Zhenyu Shan, Yuncai Liu |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2015 | Learning a Probabilistic Topology Discovering Model for Scene CategorizationabstractA recent advance in scene categorization prefers a topological based modeling to capture the existence and relationships among different scene components. To that effect, local features are typically used to handle photographing variances such as occlusions and clutters. However, in many cases, the local features alone cannot well capture the scene semantics since they are extracted from tiny regions (e.g., 4×4 patches) within an image. In this paper, we mine a discriminative topology and a low-redundant topology from the local descriptors under a probabilistic perspective, which are further integrated into a boosting framework for scene categorization. In particular, by decomposing a scene image into basic components, a graphlet model is used to describe their spatial interactions. Accordingly, scene categorization is formulated as an intergraphlet matching problem. The above procedure is further accelerated by introducing a probabilistic based representative topology selection scheme that makes the pairwise graphlet comparison trackable despite their exponentially increasing volumes. The selected graphlets are highly discriminative and independent, characterizing the topological characteristics of scene images. A weak learner is subsequently trained for each topology, which are boosted together to jointly describe the scene image. In our experiment, the visualized graphlets demonstrate that the mined topological patterns are representative to scene categories, and our proposed method beats state-of-the-art models on five popular scene data sets. Rongrong Ji, Yingjie Xia, Ying Zhang 0047, Xuelong Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2014 | A Generic Methodological Framework for Cyber-ITS: Using Cyber-infrastructure in ITS Data Analysis CasesabstractThis paper presents a method to capture the computational intensity and computing resource requirements of data analysis in intelligent transportation systems (ITS). These requirements can be transformed into a generic methodological framework for Cyber-ITS, mainly consisting of region-based ITS data divisions and tasks scheduling for processing, to support the efficient use of cyber-infrastructure (CI). To characterize the computational intensity of a particular ITS data analysis, the computational transformation is performed by data-centric and operation-centric transformation functions. The application of this framework is illustrated by two ITS data analysis cases: multi-sensor data fusion for traffic state estimation by integrating rough set theory with Dempster-Shafer (D-S) evidence theory, and geospatial computation on Global Positioning System (GPS) data for advertising value evaluation. To make the design of generic parallel computing solutions feasible for ITS data analysis in these cases, an approach is developed to decouple the region-based division from specific high performance computer (HPC) architecture and implement a prototype of the methodological framework. Experimental results show that the prototype implementation of the framework can be applied to divide the ITS data analysis into a load-balanced set of computing tasks, therefore facilitating the parallelized data fusion and geospatial computation to achieve remarkable speedup in computation time and throughput, without loss in accuracy. Yingjie Xia, Shengbao Wang |
Fundam. Informaticae | 1 |
| 2014 | Discovering anomaly on the basis of flow estimation of alert feature distributionabstractABSTRACT A challenge faced by many system administrators in utilizing the intrusion detection system (IDS) is to sift out genuine alerts buried with overwhelming alerts of benign activities generated by the IDS, especially for IDS deployed in large networks. Existing methods propose to identify the real alerts to aid the administrators. In our paper, we extend the idea of filtering irrelevant alerts based on alert volumes. And we formulate the flow estimation of abrupt changes in feature distribution caused by anomalies, by computing Kullback–Leibler distance of alert feature values under observation in comparison with a reference distribution, which is the mixture of a distribution drawing a tread from historical alerts, and a distribution derived from expertise provided by administrators. Experimental studies on the Defense Advanced Research Projects Agency dataset as well as real‐life data gathered from the IDS of a large network show that our method is able to distinguish and highlight genuine anomalies arising from the tremendous number of intrusion alerts, including different kinds of attacks and network failures. Application of this technique to alerts greatly helps the administrators in identifying real alerts and then reduces the alert load in the future. Copyright © 2013 John Wiley & Sons, Ltd. Xuejiao Liu 0002, Yingjie Xia, Yanbo Wang 0003 |
Secur. Commun. Networks | 2 |
| 2014 | Actively Learning Human Gaze Shifting Paths for Semantics-Aware Photo CroppingabstractPhoto cropping is a widely used tool in printing industry, photography, and cinematography. Conventional cropping models suffer from the following three challenges. First, the deemphasized role of semantic contents that are many times more important than low-level features in photo aesthetics. Second, the absence of a sequential ordering in the existing models. In contrast, humans look at semantically important regions sequentially when viewing a photo. Third, the difficulty of leveraging inputs from multiple users. Experience from multiple users is particularly critical in cropping as photo assessment is quite a subjective task. To address these challenges, this paper proposes semantics-aware photo cropping, which crops a photo by simulating the process of humans sequentially perceiving semantically important regions of a photo. We first project the local features (graphlets in this paper) onto the semantic space, which is constructed based on the category information of the training photos. An efficient learning algorithm is then derived to sequentially select semantically representative graphlets of a photo, and the selecting process can be interpreted by a path, which simulates humans actively perceiving semantics in a photo. Furthermore, we learn a prior distribution of such active graphlet paths from training photos that are marked as aesthetically pleasing by multiple users. The learned priors enforce the corresponding active graphlet path of a test photo to be maximally similar to those from the training photos. Experimental results show that: 1) the active graphlet path accurately predicts human gaze shifting, and thus is more indicative for photo aesthetics than conventional saliency maps and 2) the cropped photos produced by our approach outperform its competitors in both qualitative and quantitative comparisons. Yue Gao 0002, Rongrong Ji, Yingjie Xia, Qionghai Dai, Xuelong Li 0001 |
IEEE Trans. Image Process. | 4 |
| 2014 | Representative Discovery of Structure Cues for Weakly-Supervised Image SegmentationabstractWeakly-supervised image segmentation is a challenging problem with multidisciplinary applications in multimedia content analysis and beyond. It aims to segment an image by leveraging its image-level semantics (i.e., tags). This paper presents a weakly-supervised image segmentation algorithm that learns the distribution of spatially structural superpixel sets from image-level labels. More specifically, we first extract graphlets from a given image, which are small-sized graphs consisting of superpixels and encapsulating their spatial structure. Then, an efficient manifold embedding algorithm is proposed to transfer labels from training images into graphlets. It is further observed that there are numerous redundant graphlets that are not discriminative to semantic categories, which are abandoned by a graphlet selection scheme as they make no contribution to the subsequent segmentation. Thereafter, we use a Gaussian mixture model (GMM) to learn the distribution of the selected post-embedding graphlets (i.e., vectors output from the graphlet embedding). Finally, we propose an image segmentation algorithm, termed representative graphlet cut, which leverages the learned GMM prior to measure the structure homogeneity of a test image. Experimental results show that the proposed approach outperforms state-of-the-art weakly-supervised image segmentation methods, on five popular segmentation data sets. Besides, our approach performs competitively to the fully-supervised segmentation models. Yue Gao 0002, Yingjie Xia, Ke Lu 0002, Jialie Shen 0001, Rongrong Ji |
IEEE Trans. Multim. | 3 |
| 2013 | Parallelized Fusion on Multisensor Transportation Data: A Case Study in CyberITSabstractWith increasing requirements on traffic efficiency, environmental quality, energy efficiency, and economic productivity, intelligent transportation systems (ITS) research, which is centered on traffic state evaluation, meets two grand challenges. One is to process heterogeneous transportation data collected from different types of traffic sensors and the other is to reduce the high computation intensity for processing massive transportation data. To overcome both challenges, this paper proposes an approach using parallelized fusion on multisensor transportation data. Parallelized fusion is an embodied case of CyberITS framework, which is developed for the synthesis of cyberinfrastructure and ITS. The fusion functionality is shaped by a rough evidential fusion model (REFM) based on rough set and Dempster–Shafer evidence theories. The REFM consists of four components, which are sensor data input, bootstrapping rough conversion, hierarchical evidential fusion, and traffic state output. Their computation intensity is centered on conversion and fusion components, which can be optimized by the algorithm- and data-centric parallelization, respectively. Computational experiments for accuracy and efficiency demonstrate that parallelized fusion achieves a distributed, high-performance, and collaborative CyberITS implementation to provide accurate and real-time traffic state evaluation. Yingjie Xia, Xiu-mei Li, Zhenyu Shan |
Int. J. Intell. Syst. | 1 |
| 2012 | Personalized Services Recommendation Based on Context-Aware QoS PredictionabstractWith the increase of published Web services, it has become a great challenge to recommend service consumers the best services with regard to the quality of services (QoS). Collaborative filtering is often employed to predict the QoS of a specific service to a certain consumer. However, in existing collaborative filtering based service recommendation approaches, the context under which consumers submit a recommendation request is seldom taken into account when filtering similar recommenders and their corresponding experience. In this paper, we propose a new method dubbed CASR (Context-Aware Services Recommendation) by referring to previous service invocation experiences under similar context with the current consumer, which is of great importance in the personalized service recommendation system. First, the proposed algorithm clusters the service invocation records according to the similarity on context properties and selects the cluster that is most similar to the context of current consumer. Then it predicts the QoS of an unused service for current consumer based on the filtered recommendation records by Bayesian inference. Experimental results demonstrate that the proposed approach can significantly improve the accuracy of QoS prediction and service recommendation. Li Kuang, Yingjie Xia, Yuxin Mao |
ICWS | 2 |
| 2011 | A Parallel Fusion Method for Heterogeneous Multi-sensor Transportation Data
Yingjie Xia, Chengkun Wu, Qing-Jie Kong, Zhenyu Shan, Li Kuang |
MDAI | 1 |
| 2011 | Accelerating geospatial analysis on GPUs using CUDAabstractInverse distance weighting (IDW) interpolation and viewshed are two popular algorithms for geospatial analysis. IDW interpolation assigns geographical values to unknown spatial points using values from a usually scattered set of known points, and viewshed identifies the cells in a spatial raster that can be seen by observers. Although the implementations of both algorithms are available for different scales of input data, the computation for a large-scale domain requires a mass amount of cycles, which limits their usage. Due to the growing popularity of the graphics processing unit (GPU) for general purpose applications, we aim to accelerate geospatial analysis via a GPU based parallel computing approach. In this paper, we propose a generic methodological framework for geospatial analysis based on GPU and its programming model Compute Unified Device Architecture (CUDA), and explore how to map the inherent parallelism degrees of IDW interpolation and viewshed to the framework, which gives rise to a high computational throughput. The CUDA-based implementations of IDW interpolation and viewshed indicate that the architecture of GPU is suitable for parallelizing the algorithms of geospatial analysis. Experimental results show that the CUDA-based implementations running on GPU can lead to dataset dependent speedups in the range of 13–33-fold for IDW interpolation and 28–925-fold for viewshed analysis. Their computation time can be reduced by an order of magnitude compared to classical sequential versions, without losing the accuracy of interpolation and visibility judgment. Yingjie Xia, Li Kuang, Xiu-mei Li |
J. Zhejiang Univ. Sci. C | 1 |
| 2010 | Analyzing Behavioral Substitution of Web Services Based on Pi-calculusabstractThe behavioral analysis for Web services provides a priori detection of errors to ensure successful interactions in services invocation and composition, and the behavioral substitution of Web services is one of the most important issues in such analysis. In this paper, we propose to formalize the behavior of a Web service by π-calculus. Based on the formalization, we introduce two notions of behavioral substitution of Web services namely strong and weak simulation. Furthermore, we propose a derivative approach to analyzing the behavioral substitution of services according to the given notions, which is implemented based on an existing tool of π-calculus. The proposed approach takes advantage of formalization and theory of π-calculus, so that the formalized services can be naturally analyzed and the behavioral substitution of them can be easily determined. Li Kuang, Yingjie Xia, Shuiguang Deng, Jian Wu 0001 |
ICWS | 2 |
| 2006 | Portal Access to Scientific Computation and Customized Collaborative Visualization on Campus GridabstractIn order to integrate the distributed computing resources in Zhejiang University (ZJU), a campus grid is constructed by using the hierarchical combination of middleware, Globus Toolkit and Sun Grid Engine. As one realization of the computational grid, the foundation of the campus grid provides a large amount of scientific computation cycles. Furthermore, our grid system adopts two different approaches, Java 3D and ParaView, to manipulate a collaborative visualization of the computed data respectively. Therefore the users can choose the most appropriate method according to their requirements. Three cases are tested on the campus grid via portal access. The results demonstrate that the ZJU campus grid provides powerful computation cycles and an efficient collaborative visualization service Yingjie Xia, Guanghua Song |
APSCC | 1 |