EDBT 2026 Demo / reviewers in the wild / expert
Yongheng Wang
dblp:34/6716
· DBLP profile ↗
45ranked-venue papers
8as first author
34since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 3 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 12 since 2021Databases, data management, data science and information retrieval · 8 · 4 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 8 since 2021Systems, architecture and hardware · 5 · 1 first-author · 1 since 2021Computer networks · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DomainCQA: Crafting Knowledge-Intensive QA from Domain-Specific ChartsabstractChart Question Answering (CQA) evaluates Multimodal Large Language Models (MLLMs) on visual understanding and reasoning over chart data. However, existing benchmarks mostly test surface-level parsing, such as reading labels and legends, while overlooking deeper scientific reasoning. We propose DomainCQA, a framework for constructing domain-specific CQA benchmarks that emphasize both visual comprehension and knowledge-intensive reasoning. It integrates complexity-aware chart selection, multitier QA generation, and expert validation. Applied to astronomy, DomainCQA yields AstroChart, a benchmark of 1,690 QA pairs over 482 charts, exposing persistent weaknesses in fine-grained perception, numerical reasoning, and domain knowledge integration across 21 MLLMs. Fine-tuning on AstroChart improves performance across fundamental and advanced tasks. Pilot QA sets in biochemistry, economics, medicine, and social science further demonstrate DomainCQA’s generality. Together, our results establish DomainCQA as a unified pipeline for constructing and augmenting domain-specific chart reasoning benchmarks. Yujing Lu, Weiming Li 0001, Yongheng Wang, Manni Duan |
AAAI | 6 |
| 2026 | RepAttn3D: Re-parameterizing 3D attention with spatiotemporal augmentation for video understanding
Xiusheng Lu, Lechao Cheng, Sicheng Zhao, Ying Zheng 0009, Yongheng Wang, Guiguang Ding, Mingli Song |
Neural Networks | 5 |
| 2025 | AI-Driven Log Analysis: Advances and Challenges
Yongheng Wang, Wensheng Gan, Philip S. Yu |
IEEE Big Data | 1 |
| 2025 | Structure and Position-Aware Graph Modeling for Trajectory Similarity Computation Over Road NetworksabstractTrajectory similarity computation is critical to various spatial data-related applications. To date, many deep learning-based approaches have been proposed to approximate trajectory similarity. However, most of previous models focus on trajectories in Euclidean space, neglecting the information of road networks, which is an important prerequisite in many applications, such as traffic analytics, social recommendation. In this paper, we study the trajectory similarity learning over road networks. Different from previous task, trajectories over road networks contain richer and more complex information, e.g., the geographical and structure information of road networks. To this end, we propose SPGMT, a graph modeling based approach that leverages abundant structure and position information inherent in road networks for trajectory similarity learning. Particularly, our graph model learns informative node representations by simultaneously incorporating structure information of nodes from a local perspective and position information from a global perspective. This road network oriented module is the first proposal to learn from a broad context of graph topology. Afterwards, SPGMT designs a self-attention network and employs an LSTM to learn the sequential information from trajectories. We conduct experiments on real-life datasets to demonstrate the superiority of SPGMT in terms of effectiveness. Besides, additional study shows the flexibility and robustness of SPGMT. Peilun Yang, Hanchen Wang 0001, Zhangyi Xu, Zhengping Qian, Yongheng Wang, Ying Zhang 0001 |
ICDE | 5 |
| 2025 | Genetic regulation of lncRNA expression in whole human brain and their contribution to CNS disordersabstractLong non-coding RNAs (lncRNAs) play a key role in the human brain, and genetic variants regulate their expression. Herein, expression quantitative trait loci of lncRNAs encompassing 10 brain regions from 134 individuals were analyzed, and novel variants influencing lncRNA expression (eSNPs) and respective affected lncRNAs (elncRNAs) were identified. The eSNPs showed proximity to their corresponding elncRNAs, enriched in the non-coding genome, and have a high minor allele frequency. The elncRNAs exhibit a high-level and complex pattern of expression. The genetic regulation is more tissue specific for lncRNAs than for protein-coding genes, with notable differences between cerebrum and cerebellum. Nonetheless, it shows relatively similar patterns across cortex regions. Furthermore, we observed a significant enrichment of eSNPs among variants associated with neurological disorders, especially insomnia, and identified insomnia-related lncRNAs involved in immune response functions. Moreover, the present study offers an improved tool for lncRNA quantification, a novel approach for lncRNA function analysis, and a database of lncRNA expression regulation in human brain. These findings and resources will advance the research on non-coding gene expression regulation in neuroscience. Yijie He, Yaqin Tang, Pengcheng Tan, Dongyu Huang, Yongheng Wang, Tong Wen, Lizhen Shao, Qinyu Cai, Zhimou Li, Taihang Liu |
Briefings Bioinform. | 5 |
| 2025 | A unified framework for interactive visual graph matching via attribute-structure synchronization
Yuhua Liu, Jiajia Kou, Heyu Wang, Yongheng Wang, Yigang Wang, Jinchang Li, Zhiguang Zhou |
Comput. Graph. | 6 |
| 2025 | VIS4SL: A visual analytic approach for interpreting and diagnosing shortcut learning
Xiyu Meng, Tan Tang, Yuhua Zhou, Dazhen Deng, Yongheng Wang, Yingcai Wu |
Knowl. Based Syst. | 6 |
| 2025 | Dual-modality visual feature flow for medical report generationabstractMedical report generation, a cross-modal task of generating medical text information, aiming to provide professional descriptions of medical images in clinical language. Despite some methods have made progress, there are still some limitations, including insufficient focus on lesion areas, omission of internal edge features, and difficulty in aligning cross-modal data. To address these issues, we propose Dual-Modality Visual Feature Flow (DMVF) for medical report generation. Firstly, we introduce region-level features based on grid-level features to enhance the method's ability to identify lesions and key areas. Then, we enhance two types of feature flows based on their attributes to prevent the loss of key information, respectively. Finally, we align visual mappings from different visual feature with report textual embeddings through a feature fusion module to perform cross-modal learning. Extensive experiments conducted on four benchmark datasets demonstrate that our approach outperforms the state-of-the-art methods in both natural language generation and clinical efficacy metrics. Quan Tang 0006, Liming Xu, Yongheng Wang, Bochuan Zheng, Jiancheng Lv 0001, Weisheng Li 0001 |
Medical Image Anal. | 3 |
| 2025 | Deep Disease Label-guided Graph Convolutional Network for Medical Report GenerationabstractMedical report generation which extracts pathological information within medical images and subsequently produces diagnostic text autonomously aims to alleviate the workload of medical experts and offers auxiliary support in diagnoses. Despite some preliminary progress have been made, several limitations still persist, including lack of specificity in extracted visual features, insufficient consideration of cross-modal alignment and extensive preparatory work required for prior knowledge. To address these issues, we, in this article, propose a novel deep label-guided graph convolutional network for medical report generation which utilizes disease label to guide to extract pathological information from medical images. To be specific, we first construct graph convolutional network to guide the model to extract the specific visual features based on disease labels, which allowing us to selectively extract disease specificity information resided in medical images. Then, we develop cross-modal alignment module to guide the alignment across medical image, diagnose report and disease label, which enables more accurate generation with more precise description. Besides, we build pre-constructed relational matrix to guide report generation model to learn the relationship between visual features and disease types with minimal additional workload to further reduce intensive workload. Extensive experiments on three benchmark datasets, i.e., IU X-ray, MIMIC-CXR, and COV-CTR, demonstrate that the proposed method outperforms the recent state-of-the-art medical report generation methods. Ours shows a 9.2% improvement in BLEU-4 score on the IU X-ray dataset, and both BLEU-4 and CIDEr scores improve by 6.31% on the MIMIC-CXR dataset. Additionally, the results show that it can be easily to applied and extended to medical image report generation with different modalities. Liming Xu, Yongheng Wang, Quan Tang 0006, Jiancheng Lv 0001 |
ACM Trans. Knowl. Discov. Data | 2 |
| 2025 | Linking Text and Visualizations via Contextual Knowledge GraphabstractThe integration of visualizations and text is commonly found in data news, analytical reports, and interactive documents. For example, financial articles are presented along with interactive charts to show the changes in stock prices on Yahoo Finance. Visualizations enhance the perception of facts in the text while the text reveals insights of visual representation. However, effectively combining text and visualizations is challenging and tedious, which usually involves advanced programming skills. This paper proposes a semi-automatic pipeline that builds links between text and visualization. To resolve the relationship between text and visualizations, we present a method which structures a visualization and the underlying data as a contextual knowledge graph, based on which key phrases in the text are extracted, grouped, and mapped with visual elements. To support flexible customization of text-visualization links, our pipeline incorporates user knowledge to revise the links in a mixed-initiative manner. To demonstrate the usefulness and the versatility of our method, we replicate prior studies or cases in crafting interactive word-sized visualizations, annotating visualizations, and creating text-chart interactions based on a prototype system. We carry out two preliminary model tests and a user study and the results and user feedbacks suggest our method is effective. Xiwen Cai, Di Weng, Taotao Fu, Siwei Fu, Yongheng Wang, Yingcai Wu |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2025 | ChartKG: A Knowledge-Graph-Based Representation for Chart ImagesabstractChart images, such as bar charts, pie charts, and line charts, are explosively produced due to the wide usage of data visualizations. Accordingly, knowledge mining from chart images is becoming increasingly important, which can benefit downstream tasks like chart retrieval and knowledge graph completion. However, existing methods for chart knowledge mining mainly focus on converting chart images into raw data and often ignore their visual encodings and semantic meanings, which can result in information loss for many downstream tasks. In this paper, we propose ChartKG, a novel knowledge graph (KG) based representation for chart images, which can model the visual elements in a chart image and semantic relations among them including visual encodings and visual insights in a unified manner. Further, we develop a general framework to convert chart images to the proposed KG-based representation. It integrates a series of image processing techniques to identify visual elements and relations, e.g., CNNs to classify charts, yolov5 and optical character recognition to parse charts, and rule-based methods to construct graphs. We present four cases to illustrate how our knowledge-graph-based representation can model the detailed visual elements and semantic relations in charts, and further demonstrate how our approach can benefit downstream applications such as semantic-aware chart retrieval and chart question answering. We also conduct quantitative evaluations to assess the two fundamental building blocks of our chart-to-KG framework, i.e., object recognition and optical character recognition. The results provide support for the usefulness and effectiveness of ChartKG. Zhiguang Zhou, Haoxuan Wang 0001, Zhengqing Zhao, Fengling Zheng, Yongheng Wang, Wei Chen 0001, Yong Wang 0021 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2025 | ConceptThread: Visualizing Threaded Concepts in MOOC VideosabstractMassive Open Online Courses (MOOCs) platforms are becoming increasingly popular in recent years. Online learners need to watch the whole course video on MOOC platforms to learn the underlying new knowledge, which is often tedious and time-consuming due to the lack of a quick overview of the covered knowledge and their structures. In this article, we propose ConceptThread, a visual analytics approach to effectively show the concepts and the relations among them to facilitate effective online learning. Specifically, given that the majority of MOOC videos contain slides, we first leverage video processing and speech analysis techniques, including shot recognition, speech recognition and topic modeling, to extract core knowledge concepts and construct the hierarchical and temporal relations among them. Then, by using a metaphor of thread, we present a novel visualization to intuitively display the concepts based on video sequential flow, and enable learners to perform interactive visual exploration of concepts. We conducted a quantitative study, two case studies, and a user study to extensively evaluate ConceptThread. The results demonstrate the effectiveness and usability of ConceptThread in providing online learners with a quick understanding of the knowledge content of MOOC videos. Zhiguang Zhou, Lihong Cai, Lei Wang 0194, Yigang Wang, Yongheng Wang, Wei Chen 0001, Yong Wang 0021 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2024 | Improving Knowledge Distillation via Regularizing Feature Direction and Norm
Lechao Cheng, Manni Duan, Yongheng Wang, Zunlei Feng, Shu Kong |
ECCV (24) | 4 |
| 2024 | Continual Multimodal Knowledge Graph Construction
Xiang Chen 0016, Jingtian Zhang, Ningyu Zhang 0001, Tongtong Wu, Yuxiang Wang 0001, Yongheng Wang, Huajun Chen |
IJCAI | 7 |
| 2024 | HTCCN: Temporal Causal Convolutional Networks with Hawkes Process for Extrapolation Reasoning in Temporal Knowledge GraphsabstractTingxuan Chen, Jun Long, Liu Yang, Zidong Wang, Yongheng Wang, Xiongnan Jin. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Tingxuan Chen, Liu Yang 0015, Zidong Wang 0005, Yongheng Wang, Xiongnan Jin |
NAACL-HLT | 5 |
| 2024 | Deep Multiscale Fine-Grained Hashing for Remote Sensing Cross-Modal RetrievalabstractHashing retrieval is a widely used technique in high spatial resolution remote sensing (RS) images due to its efficient retrieval speed and low memory overhead. However, existing hashing retrieval methods primarily focus on matching multilabel RS images, neglecting the extensive fine-grained semantic information in cross-modal RS data. Moreover, RS images exhibit notable object size differences and contain redundant features that lack effective multiscale feature extraction methods. To address these issues, we propose a novel deep multiscale fine-grained hashing (DMFH) method for cross-modal hashing retrieval of RS data. The DMFH method comprises two modules: the feature extraction module and hashing retrieval module. In the feature extraction module, we introduce a multiscale feature representation method to extract both low-level and high-level features from RS images while using a redundant optimizer to remove duplicate features. In addition, we used embedding vectors to extract fine-grained semantic information from description texts. The hashing retrieval module uses contrastive loss and triplet loss to guide the hash function toward learning and generating hash codes from extracted features. Our proposed DMFH method achieves state-of-the-art performance in two public RS image–text datasets (RSICD and RSITMD) through extensive experiments and ablation studies. Jiaxiang Huang, Yong Feng 0002, Mingliang Zhou 0001, Xiancai Xiong, Yongheng Wang, Baohua Qiang |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2024 | Towards Effective Clustered Federated Learning: A Peer-to-Peer Framework With Adaptive Neighbor MatchingabstractIn federated learning (FL), clients may have diverse objectives, and merging all clients' knowledge into one global model will cause negative transfer to local performance. Thus, clustered FL is proposed to group similar clients into clusters and maintain several global models. In the literature, centralized clustered FL algorithms require the assumption of the number of clusters and hence are not effective enough to explore the latent relationships among clients. In this paper, without assuming the number of clusters, we propose a peer-to-peer (P2P) FL algorithm namedPANM. InPANM, clients communicate with peers to adaptively form an effective clustered topology. Specifically, we present two novel metrics for measuring client similarity and a two-stage neighbor matching algorithm based Monte Carlo method and Expectation Maximization under the Gaussian Mixture Model assumption. We have conducted theoretical analyses ofPANMon the probability of neighbor estimation and the error gap to the clustered optimum. We have also implemented extensive experiments under both synthetic and real-world clustered heterogeneity. Theoretical analysis and empirical experiments show that the proposed algorithm is superior to the P2P FL counterparts, and it achieves better performance than the centralized cluster FL method.PANMis effective even under extremely low communication budgets. Zexi Li 0001, Jiaxun Lu, Didi Zhu, Yunfeng Shao 0001, Yinchuan Li, Yongheng Wang, Chao Wu 0001 |
IEEE Trans. Big Data | 8 |
| 2024 | A Deep Spatiotemporal Trajectory Representation Learning Framework for ClusteringabstractLearning trajectory representations is essential in many Location Based Services (LBS) applications. Most traditional methods extract trajectory representations based on manually defined features, while deep learning-based methods can reduce part of the human effort. We propose a Deep Spatiotemporal Trajectory Clustering (DSTC) framework to tackle the Spatiotemporal Trajectory Representation Learning towards the Clustering-friendly space (STRLC) problem. Solving the STRLC problem is not a trivial task because: (1) Defining a uniform token size for datasets with an uneven density of trajectory data is challenging. (2) Measuring the similarity between trajectories spanning time zero in the time dimension is a problem to be solved. (3) It requires first learning a vector that can represent the overall characteristics of spatiotemporal trajectories and then mapping it to a more suitable space for clustering. To tackle these challenges, we first utilize the density-based clustering method to define tokens representing the trajectory points automatically. Then, we use polar coordinates to represent the temporal dimension of trajectories. Additionally, we improve the learned trajectory representations in a clustering-oriented latent space end to end. Experiments conducted on benchmark datasets demonstrate that DSTC achieves better accuracy than existing methods. Moreover, the representations learned from spatiotemporal trajectory data in the real world can be used to identify popular routes during the day. Yongheng Wang, Zhengxuan Lin, Xiongnan Jin, Xing Jin 0002, Di Weng, Yingcai Wu |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | Ensemble Clustering via Co-Association Matrix Self-EnhancementabstractEnsemble clustering integrates a set of base clustering results to generate a stronger one. Existing methods usually rely on a co-association (CA) matrix that measures how many times two samples are grouped into the same cluster according to the base clusterings to achieve ensemble clustering. However, when the constructed CA matrix is of low quality, the performance will degrade. In this article, we propose a simple, yet effective CA matrix self-enhancement framework that can improve the CA matrix to achieve better clustering performance. Specifically, we first extract the high-confidence (HC) information from the base clusterings to form a sparse HC matrix. By propagating the highly reliable information of the HC matrix to the CA matrix and complementing the HC matrix according to the CA matrix simultaneously, the proposed method generates an enhanced CA matrix for better clustering. Technically, the proposed model is formulated as a symmetric constrained convex optimization problem, which is efficiently solved by an alternating iterative algorithm with convergence and global optimum theoretically guaranteed. Extensive experimental comparisons with 12 state-of-the-art methods on ten benchmark datasets substantiate the effectiveness, flexibility, and efficiency of the proposed model in ensemble clustering. The codes and datasets can be downloaded at https://github.com/Siritao/EC-CMS. Yuheng Jia, Sirui Tao, Ran Wang 0001, Yongheng Wang |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | JsonCurer: Data Quality Management for JSON Based on an Aggregated SchemaabstractHigh-quality data is critical to deriving useful and reliable information. However, real-world data often contains quality issues undermining the value of the derived information. Most existing research on data quality management focuses on tabular data, leaving semi-structured data under-exploited. Due to the schema-less and hierarchical features of semi-structured data, discovering and fixing quality issues is challenging and time-consuming. To address the challenge, this paper presents JsonCurer, an interactive visualization system to assist with data quality management in the context of JSON data. To have an overview of quality issues, we first construct a taxonomy based on interviews with data practitioners and a review of 119 real-world JSON files. Then we highlight a schema visualization that presents structural information, statistical features, and quality issues of JSON data. Based on a similarity-based aggregation technique, the visualization depicts the entire JSON data with a concise tree, where summary visualizations are given above each node, and quality issues are illustrated using Bubble Sets across nodes. We evaluate the effectiveness and usability of JsonCurer with two case studies. One is in the domain of data analysis while the other concerns quality assurance in MongoDB documents. Siwei Fu, Di Weng, Yongheng Wang, Yingcai Wu |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2024 | DMA-YOLO: multi-scale object detection method with attention mechanism for aerial images
Ya-ling Li, Yong Feng 0002, Mingliang Zhou 0001, Xiancai Xiong, Yongheng Wang, Baohua Qiang |
Vis. Comput. | 5 |
| 2024 | EMPNet: An extract-map-predict neural network architecture for cross-domain recommendation
Jinpeng Chen 0001, Huan Li 0003, Hua Lu 0001, Xiongnan Jin, Kuien Liu, Yongheng Wang |
World Wide Web (WWW) | 8 |
| 2023 | Can We Edit Multimodal Large Language Models?abstractIn this paper, we focus on editing Multimodal Large Language Models (MLLMs).Compared to editing single-modal LLMs, multimodal model editing is more challenging, which demands a higher level of scrutiny and careful consideration in the editing process.To facilitate research in this area, we construct a new benchmark, dubbed MMEdit, for editing multimodal LLMs and establishing a suite of innovative metrics for evaluation.We conduct comprehensive experiments involving various model editing baselines and analyze the impact of editing different components for multimodal LLMs.Empirically, we notice that previous baselines can implement editing multimodal LLMs to some extent, but the effect is still barely satisfactory, indicating the potential difficulty of this task.We hope that our work can provide the NLP community with insights1. Siyuan Cheng 0008, Bozhong Tian, Qingbin Liu, Xi Chen 0003, Yongheng Wang, Huajun Chen, Ningyu Zhang 0001 |
EMNLP | 5 |
| 2023 | Joint Robust Representation And Generalization Enhancement For Cross-Modality Person Re-IdentificationabstractCross-modality person re-identification (cm-ReID) aims to match pedestrian images from visible and infrared cameras. Most existing methods ignore data bias due to different cameras and views and overlook the strong dependence between feature maps that hinders modal alignment. In this paper, we propose a unified method named Joint Robust Representation and Generalization Enhancement (RRGE) to alleviate the above issues. First, we propose a robust representation module (RRM), which can improve the model’s robustness for the global context, camera, and view change perturbations. Second, we propose a generalization enhancement module (GEM), which uses channel-level dropout to alleviate the dependencies between feature maps to improve the model’s generalization. Moreover, we balance the number of different modalities in each batch. Our method outperforms other state-of-the-art methods in terms of cross-modality person re-identification tasks. Heqing Cheng, Yong Feng 0002, Mingliang Zhou 0001, Xiancai Xiong, Yongheng Wang, Baohua Qiang |
ICASSP | 5 |
| 2023 | Distributed Training of Large Language ModelsabstractThe advent of large language models (LLMs), like ChatGPT ushers in revolutionary opportunities that bring a vast variety of applications (such as healthcare, law, and education) across various disciplines. The research report pointed out that the model showcases excellent performance often closely related to the parameter scale of the model, so how to train an LLM? This is a question that everyone is more concerned about. At present, there are several commonly used distributed training frameworks including Megatron-LM, DeepSpeed, etc. In this paper, we first provide a brief introduction, which refers to the current development status of LLM. Second, we start from the status, introducing the current common parallel strategies of LLM distributed training. Next, we briefly introduce the underlying technologies and frameworks that LLM relies on nowadays, describing the current popular ones and types of large models. Then we introduce the optimization techniques used in the LLMs. Finally, we summarize the problems and challenges encountered in the current LLM training and describe the possible future development direction of LLM. Fanlong Zeng, Wensheng Gan, Yongheng Wang, Philip S. Yu |
ICPADS | 3 |
| 2023 | Semantic Dissimilarity Guided Locality Preserving Projections for Partial Label Dimensionality ReductionabstractPartial label learning (PLL) is a significant weakly supervised learning framework, where each training example corresponds to a set of candidate labels among which only one is the ground-truth label. Existing works on partial label dimensionality reduction only exploit the disambiguated labels, but overlook the available semantic dissimilarity relationship hidden in the disambiguated labeling confidence, i.e., the smaller the inner product of the labeling confidences of two instances, the less likely they have the same ground-truth label. By combining such global dissimilarity relationship with local neighborhood information, we propose a novel partial label dimensionality reduction method named SDLPP, which employs an alternating procedure including candidate label disambiguation, semantic dissimilarity generation and dimensionality reduction. The labeling confidences of candidate labels and semantic dissimilarity relationship are constantly updated through the alternating procedure, where the processes in each iteration are based on the low-dimensional data obtained in the previous iteration. After the alternating procedure, SDLPP maps the original data to a pre-specified low-dimensional feature space. Comprehensive experiments on both synthetic and real-world data sets validate that SDLPP can improve the generalization performance of different PLL algorithms, and outperform state-of-the-art partial label dimensionality reduction methods. The codes can be publicly accessible on the link https://github.com/jhjiangSEU/SDLPP. Yuheng Jia, Yongheng Wang |
KDD | 3 |
| 2023 | TIAR: Text-Image-Audio Retrieval with weighted multimodal re-ranking
Peide Chi, Yong Feng 0002, Mingliang Zhou 0001, Xiancai Xiong, Yongheng Wang, Baohua Qiang |
Appl. Intell. | 5 |
| 2023 | Multi-modal transformer using two-level visual features for fake news detection
Yong Feng 0002, Xiancai Xiong, Yongheng Wang, Baohua Qiang |
Appl. Intell. | 4 |
| 2023 | Computational experiment-aided prescriptive decision-making for complex supply chains: A case of multi-generation smartphone marketing
Qingqi Long, Yingni Chen, Yongheng Wang, Shuzhu Zhang, Juanjuan Peng |
Expert Syst. Appl. | 3 |
| 2023 | Improve Session-Based Recommendation with Triplet Mining and Dynamic Perturbations Graph Neural NetworksabstractSession-based recommendation (SBR) emphasizes mining user interests to predict the next click based on recent interactions within sessions. Most current SBR methods suffer from insufficient interactive information problems and fail to distinguish session representations with high similarities, which can neglect the inherent features within sessions. To fill the gap, we propose a triplet mining enhanced graph neural networks (TME-GNN) approach to enhance the recommendation systems by mining structural and inherent information. Technically, we first generate anchor, positive and negative embeddings based on the given session and set a triplet mining task to improve the recommendation task with subtle features by pushing positive pairs close and pulling negative pairs away. Second, to robust the model, we employ a self-supervised auxiliary task by adding dynamic perturbations to the embedding space. We conduct extensive experiments to demonstrate the superiority of our method against other state-of-the-art algorithms. Our implementations are available on the following site https://github.com/Info4Rec/TME-GNN . Yong Feng 0002, Mingliang Zhou 0001, Xiancai Xiong, Yongheng Wang, Baohua Qiang, Qin Mao, Bin Fang 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 5 |
| 2023 | Satisfaction-aware Task Assignment in Spatial Crowdsourcing
Yongheng Wang, Kenli Li 0001, Xu Zhou 0001, Zhao Liu 0006, Keqin Li 0001 |
Inf. Sci. | 2 |
| 2023 | Multi-level network based on transformer encoder for fine-grained image-text matching
Yong Feng 0002, Mingliang Zhou 0001, Xiancai Xiong, Yongheng Wang, Baohua Qiang |
Multim. Syst. | 5 |
| 2023 | Multilevel Similarity-Aware Deep Metric Learning for Fine-Grained Image RetrievalabstractFast and accurate image retrieval is an important and challenging task in massive image data scenarios. As the core technology of image retrieval tasks, deep metric learning aims at learning effective embedding representations that possess two properties among data points: positive concentrated and negative separated. In this work, we propose a multilevel similarity-aware method based on deep local descriptors for deep metric learning. We take the rich interclass similarity relationship based on the deep local invariant descriptors from the data into account to optimize sampling strategies for mining informative samples. The method dynamically adjusts the margin between data points to better match the true similarity relationship between classes. Specifically, for images in a batch, we first obtain deep local descriptors and calculate the similarity matrix of the channel, pixel, and spatial levels. Then, depending on the calculated comprehensive similarity matrix, we propose a multilevel similarity-aware loss function through the deviation between pairwise distance and violate margin to make full use of informative samples. The experimental results demonstrate that our proposed method outperforms other state-of-the-art methods in terms of fine-grained image retrieval and clustering tasks. Congcong Duan, Yong Feng 0002, Mingliang Zhou 0001, Xiancai Xiong, Yongheng Wang, Baohua Qiang, Weijia Jia 0001 |
IEEE Trans. Ind. Informatics | 5 |
| 2023 | Revealing the Semantics of Data Wrangling Scripts With ComanticsabstractData workers usually seek to understand the semantics of data wrangling scripts in various scenarios, such as code debugging, reusing, and maintaining. However, the understanding is challenging for novice data workers due to the variety of programming languages, functions, and parameters. Based on the observation that differences between input and output tables highly relate to the type of data transformation, we outline a design space including 103 characteristics to describe table differences. Then, we develop Comantics, a three-step pipeline that automatically detects the semantics of data transformation scripts. The first step focuses on the detection of table differences for each line of wrangling code. Second, we incorporate a characteristic-based component and a Siamese convolutional neural network-based component for the detection of transformation types. Third, we derive the parameters of each data transformation by employing a "slot filling" strategy. We design experiments to evaluate the performance of Comantics. Further, we assess its flexibility using three example applications in different domains. Zhongsu Luo, Siwei Fu, Yongheng Wang, Mingliang Xu 0001, Yingcai Wu |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2018 | A proactive decision support method based on deep reinforcement learning and state partition
Yongheng Wang, Shaofeng Geng |
Knowl. Based Syst. | 1 |
| 2018 | Predictive complex event processing based on evolving Bayesian networks
Yongheng Wang, Guidan Chen |
Pattern Recognit. Lett. | 1 |
| 2017 | A Prediction Method Based on Complex Event Processing for Cyber Physical System
Shaofeng Geng, Xiaoxi Guo, Yongheng Wang, Renfa Li, Binghua Song |
MSN | 4 |
| 2016 | A Target-Dependent Sentiment Analysis Method for Micro-blog Streams
Yongheng Wang, Shaofeng Geng |
APWeb (2) | 1 |
| 2015 | Proactive Complex Event Processing for transportation Internet of ThingsabstractComplex Event Processing (CEP) has become the key part of Internet of Things (IoT). Proactive CEP can predict future system states and execute some actions to avoid unwanted states which brings new hope to transportation IoT. In this paper, we propose a proactive CEP architecture and method for transportation IoT. Based on basic CEP technology, this method uses structure varying Bayesian network to predict future events and system states. Different Bayesian network structures are learned and used according to different event context. A networked distributed Markov decision processes model with predicting states is proposed as sequential decision model. Q-learning method is investigated for this model to find optimal joint policy. The experimental evaluations show that this method works well when used to control congestion in transportation IoT. Yongheng Wang |
IPCCC | 1 |
| 2014 | An adaptive image processing system based on incremental learning for industrial applicationsabstractMachine learning has been applied in image processing system for object recognition, inspection and measurement. It assumes that the provided training objects are representative enough to the real objects. However in real application, new (unlearned) objects always emerge over time, which may deviate from the trained (learned) objects. The conventional image processing system using machine learning is not able to learn and then recognize these new objects. In this paper, an incremental learning based image processing system is presented. The overall system consists of three layers: execution, learning and user. The conventional image processing system is constructed in execution layer. In learning layer, adviser and incremental learning are applied to generate a new classifier. The incremental learning is differentiated into different methodologies: data accumulation and ensemble learning. Through the adviser, a proper methodology can be recommended. User is able to interact with the system via user layer. Comparing to the conventional image processing system, the proposed system is robust in industrial applications, since it deals with the classification problems dynamically. Yongheng Wang, Michael Weyrich |
ETFA | 1 |
| 2014 | A Proactive Complex Event Processing Method Based on Parallel Markov Decision Processes
Yongheng Wang, Kening Cao |
WAIM | 1 |
| 2013 | Architecture design of a vision-based intelligent system for automated disassembly of E-waste with a case study of traction batteriesabstractUnlike assembly, disassembly attracts much less attention in terms of automatization. Therefore, the paper contributes to research automation technologies for the application of disassembly. In this paper, an architecture design of intelligent vision-based system is proposed. Employing the disassembly system based on the architecture, components of electronic waste (for example, traction batteries) are detected and localized. The obtained information is then integrated in a database to determine a disassembly plan, which involves space constraint and relation between the individual components. Relying on the plan, disassembly can be implemented. The contribution of this paper is to outline the main framework of developing a vision-based intelligent disassembly cell. Michael Weyrich, Yongheng Wang |
ETFA | 2 |
| 2013 | Quality assessment of row crop plants by using a machine vision systemabstractThis paper reports research results on developing a machine vision system to assess the quality of row crop plants. Comparing to the prevalent machine vision system employed in agricultural industry for weed-crops classification as well as plant density evaluation, the proposed machine vision system is able to detect the location of plants (weed / crops) and calculate the leaves' area for plant quality assessment, even if the leaves are overlapped with each other. The developed machine vision system involves a camera system and an image processing system. The camera system uses a coaxial camera constructed by a RGB sensor and near infrared (NIR) sensor, which cooperate with a white front lighting and NIR front lighting respectively. Plants are firstly captured by the coaxial camera. The plants are segmented from background on RGB image; the overlapping edges of leaves are detected on NIR image. Afterwards the overlapping leaves are separated and assigned to the assessed stem position of plants. At last, based on the assigned leaves, the plants are separated, and the area of plant canopy is calculated. A set of experiments have been made to prove the feasibility of the proposed machine vision system. Michael Weyrich, Yongheng Wang, Matthias Scharf |
IECON | 2 |
| 2011 | Mechatronic engineering of novel manufacturing processes implemented by modular and sensor-guided machineryabstractNovel manufacturing processes and machine designs can be developed by means of mechatronic modules and their adaptation to the specific manufacturing case. The complexity of these modules must be carefully chosen so that a reuse for different application processes becomes possible. If a mechatronic module is too specific or complex, it cannot be engaged for other machine designs, as it is made only for that specific case. The methodology of axiomatic design is introduced and adapted towards the engineering of mechatronic modules in machine development for special manufacturing processes. The presented approach of systematic modular design allows for the engineering of mechatronic modules and the identification of the required links between modules. Michael Weyrich, Philipp Klein, Martin Laurowski, Yongheng Wang |
ETFA | 4 |
| 2005 | Parallel Mining of Top-K Frequent Itemsets in Very Large Text Database
Yongheng Wang, Yan Jia 0001, Shuqiang Yang |
WAIM | 1 |