Jingru Yang

dblp:227/9414 · DBLP profile ↗
← Back
24ranked-venue papers
8as first author
18since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Databases, data management, data science and information retrieval · 7 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Procore: Robust Core-Set Selection Via Pareto Multi-Dimensional Optimization From Noisy Data
Xiaoou Ding, Hongbin Hu, Songnan Jiang, Muyun Zhou, Chen Wang 0018, Jingru Yang, Hongzhi Wang 0001
ICDE6
2026 GasSeg: A lightweight real-time infrared gas segmentation network for edge devices
Huan Yu 0002, Jin Wang 0015, Jingru Yang, Kaixiang Huang, Fengtao Deng, Guodong Lu, Shengfeng He
Pattern Recognit.3
2026 An Intelligent Multitask Framework for Industrial Gas Leak Detection and Analysis With Infrared Optical Gas Imaging
abstract
Infrared (IR) optical gas imaging (OGI) is widely adopted in industrial environments for detecting fugitive gas emissions. However, conventional IR OGI systems rely heavily on manual inspection, lacking capabilities for active leak localization and in-depth analysis, which increases labor costs and risks of human error. To address these challenges, we present LeakHunter, an intelligent multitask framework designed for industrial gas leak monitoring and decision support. LeakHunter integrates seamlessly with IR cameras and can be deployed on edge computing devices, enabling real-time, on-site leak detection in harsh industrial settings. At the core of LeakHunter is a novel keypoint detection paradigm tailored for IR OGI, capable of localizing both leak sources and diffusion endpoints to enable effective spatiotemporal trend analysis. The framework also estimates critical leak attributes, including plume morphology and flow rate, supporting rapid and informed response. To further enhance detection accuracy, we introduce a biomimetic attention module that improves gas-background separation under complex thermal conditions, and a collaborative multitask head for efficient cross-task feature sharing. In addition, two benchmark datasets are proposed, one of which is a field-test set collected in real industrial scenarios. Experiments demonstrate that LeakHunter achieves state-of-the-art performance across multiple tasks, with an F2 score of 94.8% for gas segmentation and 97.9% for leak keypoint localization, while running at 28.4 FPS on a portable IR OGI device. These results highlight its potential as a deployable, intelligent solution for enhancing industrial safety and automation.
Huan Yu 0002, Jin Wang 0015, Jingru Yang, Kaixiang Huang, Fengtao Deng, Zaixing He, Guodong Lu
IEEE Trans. Ind. Informatics3
2026 Purified Zero-Shot Sketch-Based Image Retrieval
abstract
Sketches, as a new solution in multimedia systems that can replace natural language, are characterized by sparse visual cues such as simple strokes that differ significantly from natural images containing complex elements such as background, foreground, and texture. This misalignment poses substantial challenges for zero-shot sketch-based image retrieval (ZS-SBIR). Prior approaches match sketches to full images and tend to overlook redundant elements in natural images, leading to model distraction and semantic ambiguity. To address this issue, we introduce a distraction-agnostic framework, purified cross-domain matching (PuXIM), which operates on a straightforward principle: masking and matching. We devise a visual-cross-linguistic (VxL) sampler that generates linguistic masks based on semantic labels to obscure semantically irrelevant image features. Our novel contribution is the concept of purified masked matching (PMM), which comprises two processes: (1)reconstruction, which compels the image encoder to reconstruct the masked image feature, and (2)interaction, which involves a transformer decoder that processes both sketch and masked image features to investigate cross-domain relationships for effective matching. Evaluated on the TU-Berlin, Sketchy, and QuickDraw datasets, PuXIM sets new benchmarks in terms of performance. Importantly, the distraction-agnostic nature of the matching process renders PuXIM more conducive to training, enabling efficient adaptation to zero-shot scenarios with reduced data requirements and low data quality.
Jingru Yang, Jin Wang 0015, Kaixiang Huang, Guodong Lu, Shengfeng He
IEEE Trans. Multim.2
2026 GranSSG: Correlating Volumetric Granularities for 3D Semantic Scene Graph Prediction
abstract
Predicting 3D Semantic Scene Graphs (3DSSG) is vital for understanding complex scenes by constructing structured representations. Current methods struggle with significant granularity discrepancies among instances, often relying on features at a single scale, which hampers their ability to perceive and interact with differently sized instances. To tackle this challenge, we introduce GranSSG, a novel approach that integrates volumetric granular awareness into 3DSSG prediction. Central to GranSSG is the Volumetric Pooling block, which aggregates features from multiple instance volumes, enhancing the representation of instance patterns across different granularities. Complementing this, the Granularity Transformer block dynamically directs attention to instance features across various network layers, ensuring precise perception of instances regardless of their granularity. Furthermore, the Cross-Granularity Correlation Transformer block mitigates performance degradation in instance pair relationship prediction by adaptively fusing hybrid features from different granularities, providing a comprehensive representation of instance pairs. Extensive evaluations on the challenging 3DSSG benchmark demonstrate that GranSSG significantly enhances prediction performance, setting a new state-of-the-art in 3DSSG prediction.
Kaixiang Huang, Jin Wang 0015, Jingru Yang, Jiao Yi, Guodong Lu, Shengfeng He
IEEE Trans. Vis. Comput. Graph.4
2025 LIFTus: An Adaptive Multi-Aspect Column Representation Learning for Table Union Search
abstract
Table union search (TUS) represents a fundamental operation in data lakes to find tables unionable to the given one. Recent approaches to TUS mainly learn column representations for searching by introducing Pre-trained Language Models (PLMs), especially on columns with linguistic data. However, a significant amount of non-linguistic data, notably represented by domain-specific strings and numerical data in the data lake, are still under-explored in the existing methods. To address this issue, we propose LIFTus, an adaptive multi-aspect column representation for table unionable search, where aspect refers to a concept more flexible than data types, so that a single column can exhibit multiple aspects simultaneously. LIFTus aims at combining different aspects of a column (including both linguistic and non-linguistic aspects) to promote the effectiveness and generalization of TUS in a self-supervised manner. Specifically, besides employing PLMs to extract the linguistic aspects from an individual column, LIFTus trains a pattern encoder to learn possible character-level sequential patterns for the column, and builds a number encoder to capture numerical aspects of the column, including the distribution and magnitude features. LIFTus further utilizes a hierarchical cross-attention aided by aspect-relevant statistics to combine these aspects adaptively in producing the final column representations, which are indexed by vector retrieval techniques to achieve efficient search. Extensive experimental results demonstrate that LIFTus has outperformed the current state-of-the-art methods in terms of effectiveness, and achieved much better generalization capability to support unseen data.
Ermu Qiu, Jun Gao 0003, Yaofeng Tu, Jingru Yang
ICDE4
2025 ServerlessIE: A Reusable Information Extraction System with Serverless Function
abstract
Information extraction (IE) plays a pivotal role for data-driven applications. However, existing IE systems face challenges in poor generalization and re-usability. They perform well within specific domains or tasks but struggle to generalize to new domains or tasks. Meanwhile, it is difficult for users to reuse or replicate existing information extraction algorithms. In this paper, we present ServerlessIE, a reusable IE framework supporting multi-modal data (text, images, videos) via serverless functions. It unifies heterogeneous extraction algorithms and simplifies their replication across tasks. We also develop a deep learning-based function recommendation model to help users find appropriate functions for their tasks more conveniently. Experiments demonstrate the effectiveness of our approach in cross-modal extraction and adaptive function recommendation.
Jingru Yang, Lingran Bu, Jinfeng Wen, Yi Liu 0014
ICWS1
2025 Art4Math: Handwritten Mathematical Expression Recognition via Multimodal Sketch Grounding
Jin Wang 0015, Kaixiang Huang, Guodong Lu, Jingru Yang, Shengfeng He
ACM Multimedia6
2025 Jury-and-Judge Chain-of-Thought for Uncovering Toxic Data in 3D Visual Grounding
abstract
3D Visual Grounding (3DVG) faces persistent challenges due to coarse scene-level observations and logically inconsistent annotations, which introduce ambiguities that compromise data quality and hinder effective model supervision. To address these challenges, we introduce Refer-Judge, a novel framework that harnesses the reasoning capabilities of Multimodal Large Language Models (MLLMs) to identify and mitigate toxic data. At the core of Refer-Judge is a Jury-and-Judge Chain-of-Thought paradigm, inspired by the deliberative process of the judicial system. This framework targets the root causes of annotation noise: jurors collaboratively assess 3DVG samples from diverse perspectives, providing structured, multi-faceted evaluations. Judges then consolidate these insights using a Corroborative Refinement strategy, which adaptively reorganizes information to correct ambiguities arising from biased or incomplete observations. Through this two-stage deliberation, Refer-Judge significantly enhances the reliability of data judgments. Extensive experiments demonstrate that our framework not only achieves human-level discrimination at the scene level but also improves the performance of baseline algorithms via data purification. Code is available at https://github.com/Hermione-HKX/Refer_Judge.
Kaixiang Huang, Jin Wang 0015, Jingru Yang, Huan Yu 0002, Guodong Lu, Shengfeng He
NeurIPS4
2025 ExtRep: a GUI test repair method for mobile applications based on test-extension
Chu Zeng, Xiangping Chen, Xing Chen 0002, Xiaocong Zhou, Jingru Yang, Gang Huang 0001, Zibin Zheng
Autom. Softw. Eng.7
2025 A lightweight and robust detection network for diverse glass surface defects via scale- and shape-aware feature extraction
Huan Yu 0002, Jin Wang 0015, Jingru Yang, Yiming Liang, Zhan Wang 0002, Haiyan He, Guodong Lu
Eng. Appl. Artif. Intell.3
2025 Sketch-SparseNet: Sparse convolution framework for sketch recognition
Jingru Yang, Jin Wang 0015, Guodong Lu, Huan Yu 0002, Heming Fang, Shengfeng He
Pattern Recognit.1
2024 MeDiC: Metasearch Service on Distributed Confidential Data
abstract
Traditional search engines aggregate vast amount of data on the Internet to provide keyword search services. However, in some privacy-sensitive fields like healthcare and e-government, the personal data often contains a substantial amount of private or sensitive information across different stakeholders’ data repositories. Due to the presence of private or sensitive information within the raw data or its metadata, as well as the lack of unified local search models, employing the approach with global data aggregation and indexing is not feasible. To solve this problem, we propose MeDiC, a metasearch service for discovering globally distributed confidential data. By distributing search requests through a unified interface, MeDiC offers a metasearch service to users without aggregating distributed confidential data. Experiments demonstrate that the MeDiC exhibits strong extensibility and user friendliness to enable developers to select different models, algorithms and search parameters based on specific search scenarios.
Wenchun Jing, Jingru Yang, Yun Ma 0002, Yi Liu 0014, Chaoran Luo
ICWS2
2024 Cross-Modal Pixel-and-Stroke representation aligning networks for free-hand sketch recognition
Jin Wang 0015, Jingru Yang, Ping Ni, Guodong Lu, Heming Fang, Huan Yu 0002, Kaixiang Huang
Expert Syst. Appl.3
2024 IEFM and IDS: Enhancing 3D environment perception via information encoding in indoor point cloud semantic segmentation
Kaixiang Huang, Jin Wang 0015, Jingru Yang, Guodong Lu, Huan Yu 0002
Neurocomputing3
2024 MsVFE and V-SIAM: Attention-based multi-scale feature interaction and fusion for outdoor LiDAR semantic segmentation
Jingru Yang, Jin Wang 0015, Kaixiang Huang, Guodong Lu, Huan Yu 0002, Wenming Zou
Neurocomputing1
2024 Granular3D: Delving into multi-granularity 3D scene graph prediction
abstract
This paper addresses the significant challenges in 3D Semantic Scene Graph (3DSSG) prediction, essential for understanding complex 3D environments. Traditional approaches, primarily using PointNet and Graph Convolutional Networks , struggle with effectively extracting multi-grained features from intricate 3D scenes , largely due to a focus on global scene processing and single-scale feature extraction. To overcome these limitations, we introduce Granular3D, a novel approach that shifts the focus towards multi-granularity analysis by predicting relation triplets from specific sub-scenes. One key is the Adaptive Instance Enveloping Method (AIEM), which establishes an approximate envelope structure around irregular instances, providing shape-adaptive local point cloud sampling, thereby comprehensively covering the contextual environments of instances. Moreover, Granular3D incorporates a Hierarchical Dual-Stage Network (HDSN), which differentiates and processes features of instances and their pairs at varying scales, leading to a targeted prediction of instance categories and their relationships. To advance the perception of sub-scene in HDSN, we design a Gather Point Transformer structure (GaPT) that enables the combinatorial interaction of local information from multiple point cloud sets, achieving a more comprehensive local contextual feature extraction. Extensive evaluations on the challenging 3DSSG benchmark demonstrate that our methods provide substantial improvements, establishing a new state-of-the-art in 3DSSG prediction, boosting the top-50 triplet accuracy by +2.8%.
Kaixiang Huang, Jingru Yang, Jin Wang 0015, Shengfeng He, Zhan Wang 0002, Haiyan He, Guodong Lu
Pattern Recognit.2
2021 A Human-in-the-loop Approach to Social Behavioral Targeting
abstract
Behavioral targeting plays an important role in social media advertising for capturing users' preferences of ads. While existing studies of behavioral targeting mainly focus on the user behaviors that have explicit correlations with ads, such as ad clicking and web search, many implicit relationships between users and ads, which reside in a variety of heterogeneous sources in social media platforms, are not utilized to enhance the prediction of users' preferences of ads.In this paper, we propose a two-pronged approach to behavioral targeting that effectively addresses the above difficulties. First, we model the implicit relationships between users and ads as a heterogeneous information network (HIN), and propose a method that first performs representation learning in the HIN and then uses the learned representations to train a prediction model for boosting the performance of behavioral targeting. Second, we develop a human-in-the-loop framework to address the incompleteness challenge in HIN construction that may result in inferior performance of model prediction. The framework judiciously selects the most "beneficial" tasks to ask human for completing the HIN and utilizes the results from human to update the representation learning of HIN. We validate the effectiveness of our approach through extensive experiments on real datasets collected from WeChat, the largest social media platform in China. The experimental results show that our approach is effective at constructing a high-quality HIN at a low cost of human involvement, and the HIN can significantly improve the performance of social behavioral targeting.
Jingru Yang, Xiaoman Zhao, Ju Fan, Xiaoyong Du 0001
ICDE1
2020 A game-based framework for crowdsourced data labeling
Jingru Yang, Ju Fan, Zhewei Wei, Guoliang Li 0001, Tongyu Liu, Xiaoyong Du 0001
VLDB J.1
2019 CrowdGame: A Game-Based Crowdsourcing System for Cost-Effective Data Labeling
abstract
Large-scale data labeling has become a major bottleneck for many applications, such as machine learning and data integration. This paper presents CrowdGame, a crowdsourcing system that harnesses the crowd to gather data labels in a cost-effective way. CrowdGame focuses on generating high-quality labeling rules to largely reduce the labeling cost while preserving quality. It first generates candidate rules, and then devises a game-based crowdsourcing approach to select rules with high coverage and accuracy. CrowdGame applies the generated rules for effective data labeling. We have implemented CrowdGame and provided a user-friendly interface for users to deploy their labeling applications. We will demonstrate CrowdGame in two representative data labeling scenarios, entity matching and relation extraction.
Tongyu Liu, Jingru Yang, Ju Fan, Zhewei Wei, Guoliang Li 0001, Xiaoyong Du 0001
SIGMOD Conference2
2019 Binary Image Carving for 3D Printing
Jingru Yang, Sha He, Lin Lu 0001
Comput. Aided Des.1
2019 3D printed perforated QR codes
Jingru Yang, Hao Peng 0001, Lin Lu 0001
Comput. Graph.1
2019 Distribution-Aware Crowdsourced Entity Collection
abstract
The problem of crowdsourced entity collection solicits people (a.k.a. workers) to complete missing data in a database and has witnessed many applications in knowledge base completion and enterprise data collection. Although previous studies have attempted to address the “open world” challenge of crowdsourced entity collection, they do not pay much attention to the “distribution” of the collected entities. Evidently, in many real applications, users may have distribution requirements on the collected entities, e.g., even spatial distribution when collecting points-of-interest. In this paper, we study a new research problem, distribution-aware crowdsourced entity collection (CrowdDEC): Given an expected distribution w.r.t. an attribute (e.g., region or year), it aims to collect a set of entities via crowdsourcing and minimize the difference of the entity distribution from the expected distribution. Due to the openness of crowdsourcing, the CrowdDEC problem calls for effective crowdsourcing quality control. We propose an adaptive worker selection approach to address this problem. The approach estimates underlying entity distribution of workers on-the-fly based on the collected entities. Then, it adaptively selects the best set of workers that minimizes the difference from the expected distribution. Once workers submit their answers, it adjusts the estimation of workers' underlying distributions for subsequent adaptive worker selections. We prove the hardness of the problem, and develop effective estimation techniques as well as efficient worker selection algorithms to support this approach. We deployed the proposed approach on Amazon Mechanical Turk and the experimental results on two real datasets show that the approach achieves superiority on both effectiveness and efficiency.
Ju Fan, Zhewei Wei, Dongxiang Zhang, Jingru Yang, Xiaoyong Du 0001
IEEE Trans. Knowl. Data Eng.4
2018 Cost-Effective Data Annotation using Game-Based Crowdsourcing
abstract
Large-scale data annotation is indispensable for many applications, such as machine learning and data integration. However, existing annotation solutions either incur expensive cost for large datasets or produce noisy results. This paper introduces a cost-effective annotation approach, and focuses on the labeling rule generation problem that aims to generate high-quality rules to largely reduce the labeling cost while preserving quality. To address the problem, we first generate candidate rules, and then devise a game-based crowdsourcing approach C ROWD G AME to select high-quality rules by considering coverage and precision. C ROWD G AME employs two groups of crowd workers: one group answers rule validation tasks (whether a rule is valid) to play a role of rule generator, while the other group answers tuple checking tasks (whether the annotated label of a data tuple is correct) to play a role of rule refuter. We let the two groups play a two-player game: rule generator identifies high-quality rules with large coverage and precision, while rule refuter tries to refute its opponent rule generator by checking some tuples that provide enough evidence to reject rules covering the tuples. This paper studies the challenges in C ROWD G AME . The first is to balance the trade-off between coverage and precision. We define the loss of a rule by considering the two factors. The second is rule precision estimation. We utilize Bayesian estimation to combine both rule validation and tuple checking tasks. The third is to select crowdsourcing tasks to fulfill the game-based framework for minimizing the loss. We introduce a minimax strategy and develop efficient task selection algorithms. We conduct experiments on entity matching and relation extraction, and the results show that our method outperforms state-of-the-art solutions.
Jingru Yang, Ju Fan, Zhewei Wei, Guoliang Li 0001, Tongyu Liu, Xiaoyong Du 0001
Proc. VLDB Endow.1