Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Zongyu Wang

dblp:124/7848 · DBLP profile ↗
← Back
13ranked-venue papers
1as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Information extraction and text analysis · 41% Language models and text generation · 26% Knowledge representation and reasoning · 16%
Software engineering, system software, and programming languages
1 paper
Program synthesis and code generation · 50% Software maintenance and evolution · 50%
Computer graphics and multimedia
1 paper
Multimedia systems and quality of experience · 77% Visual content generation and editing · 23%
Databases, data mining, and information retrieval
2 papers
Knowledge graphs · 100%

Topics — the 15 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Software maintenance and evolution › program comprehension
code comprehension
1.012026
CoreCodeBench: Decoupling Code Intelligence via Fine-Grained Repository-Level Tasks · ACL (1) 2026
Natural language and speech › Language models and text generation
large language model reasoning
0.912025
I2CR: Intra- and Inter-modal Collaborative Reflections for Multimodal Entity Linking · ACM Multimedia 2025
Natural language and speech › Information extraction and text analysis › entity linking
multimodal entity linking
0.912025
I2CR: Intra- and Inter-modal Collaborative Reflections for Multimodal Entity Linking · ACM Multimedia 2025
Multimedia systems and quality of experience
visual quality assessment
0.912025
Q-Eval-100K: Evaluating Visual Quality and Alignment Level for Text-to-Vision Content · CVPR 2025
Natural language and speech › Information extraction and text analysis › lexical semantics › multiword expression
noun compound interpretation
0.712023
Noun Compound Interpretation With Relation Classification and Paraphrasing · IEEE Trans. Knowl. Data Eng. 2023
Natural language and speech › Language models and text generation › text generation
paraphrase generation
0.712023
Noun Compound Interpretation With Relation Classification and Paraphrasing · IEEE Trans. Knowl. Data Eng. 2023
Natural language and speech › Information extraction and text analysis › relation extraction
relation classification
0.712023
Noun Compound Interpretation With Relation Classification and Paraphrasing · IEEE Trans. Knowl. Data Eng. 2023
Natural language and speech › Information extraction and text analysis
slot filling
0.712023
Noun Compound Interpretation With Relation Classification and Paraphrasing · IEEE Trans. Knowl. Data Eng. 2023
Knowledge graphs
taxonomy expansion
0.712023
Towards Visual Taxonomy Expansion · ACM Multimedia 2023
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge graph
entity alignment
0.612022
An Effective and Efficient Entity Alignment Decoding Algorithm via Third-Order Tensor Isomorphism · ACL (1) 2022
Knowledge, reasoning and agents › Knowledge representation and reasoning
knowledge graph
0.612022
An Effective and Efficient Entity Alignment Decoding Algorithm via Third-Order Tensor Isomorphism · ACL (1) 2022
Natural language and speech › Language models and text generation
code language models
0.312026
CoreCodeBench: Decoupling Code Intelligence via Fine-Grained Repository-Level Tasks · ACL (1) 2026
Computer vision › Vision and language
multimodal reasoning
0.312025
I2CR: Intra- and Inter-modal Collaborative Reflections for Multimodal Entity Linking · ACM Multimedia 2025
Visual content generation and editing › image generation
text-to-image generation
0.312025
Q-Eval-100K: Evaluating Visual Quality and Alignment Level for Text-to-Vision Content · CVPR 2025
Knowledge graphs
link prediction
0.212022
An Effective and Efficient Entity Alignment Decoding Algorithm via Third-Order Tensor Isomorphism · ACL (1) 2022

Methods — techniques the papers use, named apart from their topics

large language model · 2.0mean opinion score · 1.7third-order tensor isomorphism · 1.1graph neural network · 1.1multi-round iterative reasoning · 0.9intra- and inter-modal collaborative reflection · 0.9context prompts · 0.9context prompt · 0.9visual prototype learning · 0.7textual hypernymy learning · 0.7multi-view representation learning · 0.7hyper-proto constraint · 0.7contrastive learning · 0.7
YearPublicationVenuePosition
2026 CoreCodeBench: Decoupling Code Intelligence via Fine-Grained Repository-Level Tasks
abstract
Lingyue Fu, Hao Guan, Bolun Zhang, Haowei Yuan, Yaoming Zhu, Lin Qiu, ZongYu Wang, Xuezhi Cao, Xunliang Cai, Weiwen Liu, Weinan Zhang, Yong Yu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Lingyue Fu, Hao Guan 0001, Haowei Yuan, Yaoming Zhu, Zongyu Wang, Xuezhi Cao, Weiwen Liu, Weinan Zhang 0001, Yong Yu 0001
ACL (1)7
2026 Robust unsupervised visual tracking via image-to-video identity knowledge transferring
Bin Kang, Zongyu Wang, Dong Liang 0008, Tianyu Ding, Songlin Du
Pattern Recognit.2
2025 Q-Eval-100K: Evaluating Visual Quality and Alignment Level for Text-to-Vision Content
abstract
Evaluating text-to-vision content hinges on two crucial aspects: visual quality and alignment. While significant progress has been made in developing objective models to assess these dimensions, the performance of such models heavily relies on the scale and quality of human annotations. According to Scaling Law, increasing the number of human-labeled instances follows a predictable pattern that enhances the performance of evaluation models. Therefore, we introduce a comprehensive dataset designed to Evaluate Visual quality and Alignment Level for text-to-vision content (Q-EVAL-100K), featuring the largest collection of human-labeled Mean Opinion Scores (MOS) for the mentioned two aspects. The Q-EVAL-100K dataset encompasses both text-to-image and text-to-video models, with 960K human annotations specifically focused on visual quality and alignment for 100K instances (60K images and 40K videos). Leveraging this dataset with context prompt, we propose Q-Eval-Score, a unified model capable of evaluating both visual quality and alignment with special improvements for handling long-text prompt alignment. Experimental results indicate that the proposed Q-Eval-Score achieves superior performance on both visual quality and alignment, with strong generalization capabilities across other benchmarks. These findings highlight the significant value of the Q-EVAL-100K dataset. Data and codes will be available at https://github.com/zzc-1998/Q-Eval.
Tengchuan Kou, Shushi Wang, Chunyi Li 0001, Wei Sun 0029, Wei Wang 0213, Zongyu Wang, Xuezhi Cao, Xiongkuo Min, Xiaohong Liu 0001, Guangtao Zhai
CVPR8
2025 I2CR: Intra- and Inter-modal Collaborative Reflections for Multimodal Entity Linking
abstract
Multimodal entity linking plays a crucial role in a wide range of applications. Recent advances in large language model-based methods have become the dominant paradigm for this task, effectively leveraging both textual and visual modalities to enhance performance. Despite their success, these methods still face two challenges, including unnecessary incorporation of image data in certain scenarios and the reliance only on a one-time extraction of visual features, which can undermine their effectiveness and accuracy. To address these challenges, we propose a novel LLM-based framework for the multimodal entity linking task, called Intra- and Inter-modal Collaborative Reflections. This framework prioritizes leveraging text information to address the task. When text alone is insufficient to link the correct entity through intra- and inter-modality evaluations, it employs a multi-round iterative strategy that integrates key visual clues from various aspects of the image to support reasoning and enhance matching accuracy. Extensive experiments on three widely used public datasets demonstrate that our framework consistently outperforms current state-of-the-art methods in the task, achieving improvements of 3.2%, 5.1%, and 1.6%, respectively. Our code is available at https://github.com/ziyan-xiaoyu/I2CR/.
Junwen Li, Tong Ruan, Chao Wang 0095, Xinyan He, Zongyu Wang, Xuezhi Cao
ACM Multimedia7
2025 Progressive Masking Oriented Self-Taught Learning for Occluded Facial Expression Recognition
abstract
Self-taught learning (STL) is a promising solution that reduces the performance gap between weakly supervised and fully supervised learning for easily accessible, label-free images. The success of traditional STL solutions relies on the assumption that the target appearance is completely visible and well-defined. In real-world facial expression recognition scenarios, however, saliency regions are often partially occluded, which significantly hampers the generalization capability of STL methods. Nevertheless, few studies have investigated the impact of occlusion on STL. In this paper, we propose an interweaved autoencoder network for weakly supervised facial expression recognition in occlusion scenarios. The key innovation of our network lies in the Residual Connection Union (RCU) blocks that can integrate the Convolutional Neural Network (CNN) and Transformer layers into a multi-scale structure. The RCU enables a progressive masking strategy to accurately identify and focus on contributive yet often overlooked image patches by analyzing the relationships among region-level target representations. In addition, we introduce a self-knowledge distillation module for the effective training of the proposed autoencoder network. Extensive experiments are conducted on four public datasets to demonstrate the superiority of our method over related works.
Bin Kang, Shuangshuang Wang, Zongyu Wang, Haie Dou, Lei Wang 0009, Zhijie Xia
IEEE Trans. Affect. Comput.3
2024 Negation Triplet Extraction with Syntactic Dependency and Semantic Consistency
abstract
Previous works of negation understanding mainly focus on negation cue detection and scope resolution, without identifying negation subject which is also significant to the downstream tasks. In this paper, we propose a new negation triplet extraction (NTE) task which aims to extract negation subject along with negation cue and scope. To achieve NTE, we devise a novel Syntax&Semantic-Enhanced Negation Extraction model, namely SSENE, which is built based on a generative pretrained language model (PLM) of Encoder-Decoder architecture with a multi-task learning framework. Specifically, the given sentence’s syntactic dependency tree is incorporated into the PLM’s encoder to discover the correlations between the negation subject, cue and scope. Moreover, the semantic consistency between the sentence and the extracted triplet is ensured by an auxiliary task learning. Furthermore, we have constructed a high-quality Chinese dataset NegComment based on the users’ reviews from the real-world platform of Meituan, upon which our evaluations show that SSENE achieves the best NTE performance compared to the baselines. Our ablation and case studies also demonstrate that incorporating the syntactic information helps the PLM’s recognize the distant dependency between the subject and cue, and the auxiliary task learning is helpful to extract the negation triplets with more semantic consistency. We further demonstrate that SSENE is also competitive on the traditional CDSR task.
Deqing Yang, Yanghua Xiao, Zongyu Wang
LREC/COLING5
2024 Fundamental Limits of Direction Finding in Distributed Arrays Exploiting Auxiliary Sources
abstract
We consider the problem of estimating the directions of multiple target sources by exploiting auxiliary sources, focusing on a single snapshot obtained by the distributed array with position errors and angular offsets of subarrays. Former calibration methods generally assume the directions of auxiliary sources are unknown, while prior knowledge of the auxiliary sources is usually available in practice. In order to quantify the effects of auxiliary sources and their prior information, we model the directions of auxiliary sources as Gaussian random variables and use their standard deviations to quantify the prior information. We derive the prior Cramér-Rao lower bound (CRB) of the direction estimations in the new model. Simulation results show that calibration with auxiliary sources performs better than self-calibration and the prior CRB is much lower than the existing counterparts assuming unknown auxiliary sources, implying much potential to improve the estimation performance by employing the prior information.
Zongyu Wang, Yuhan Li 0006, Yihan Su, Tianyao Huang, Yimin Liu 0003
ICASSP1
2024 Exploiting Duality in Open Information Extraction with Predicate Prompt
abstract
Open information extraction (OpenIE) aims to extract the schema-free triplets in the form of (subject, predicate, object) from a given sentence. Compared with general information extraction (IE), OpenIE poses more challenges for the IE models, especially when multiple complicated triplets exist in a sentence. To extract these complicated triplets more effectively, in this paper we propose a novel generative OpenIE model, namely DualOIE, which achieves a dual task at the same time as extracting some triplets from the sentence, i.e., converting the triplets into the sentence. Such dual task encourages the model to correctly recognize the structure of the given sentence and thus is helpful to extract all potential triplets from the sentence. Specifically, DualOIE extracts the triplets in two steps: 1) first extracting a sequence of all potential predicates, 2) then using the predicate sequence as a prompt to induce the generation of triplets. Our experiments on two benchmarks and our dataset constructed from Meituan demonstrate that DualOIE achieves the best performance among the state-of-the-art baselines. Furthermore, the online A/B test on Meituan platform shows that 0.93% improvement of QV-CTR and 0.56% improvement of UV-CTR have been obtained when the triplets extracted by DualOIE were leveraged in Meituan's search system.
Zhen Chen 0035, Deqing Yang, Yanghua Xiao, Zongyu Wang, Rui Xie 0005, Yunsen Xian
WSDM6
2023 Towards Visual Taxonomy Expansion
abstract
Taxonomy expansion task is essential in organizing the ever-increasing volume of new concepts into existing taxonomies. Most existing methods focus exclusively on using textual semantics, leading to an inability to generalize to unseen terms and the "Prototypical Hypernym Problem." In this paper, we propose Visual Taxonomy Expansion (VTE), introducing visual features into the taxonomy expansion task. We propose a textual hypernymy learning task and a visual prototype learning task to cluster textual and visual semantics. In addition to the tasks on respective modalities, we introduce a hyper-proto constraint that integrates textual and visual semantics to produce fine-grained visual semantics. Our method is evaluated on two datasets, where we obtain compelling results. Specifically, on the Chinese taxonomy dataset, our method significantly improves accuracy by 8.75%. Additionally, our approach performs better than ChatGPT on the Chinese taxonomy dataset.
Tinghui Zhu, Jiaqing Liang, Haiyun Jiang, Yanghua Xiao, Zongyu Wang, Rui Xie 0005, Yunsen Xian
ACM Multimedia6
2023 Noun Compound Interpretation With Relation Classification and Paraphrasing
abstract
Noun compounds are abundant in various languages and their interpretations have been applied in a wide range of NLP tasks. However, most existing work only uses relation classification- or paraphrasing-based methods to model this problem, failing in coverage or accuracy. We argue that the above two approaches are complementary to each other for the noun compound interpretation. In this paper, we propose a two-phase strategy to solve this task. The first phase is to perform the relation classification sub-task with a novel multi-view representation learning model. When noun compounds are predicted as the non-semantic relation, i.e., NA, or the confidence scores are below the threshold, the second phase, namely paraphrasing, will be triggered to interpret noun compounds with a contrastive slot filling method. To evaluate the effectiveness of our methods, we construct the largest Chinese dataset for noun compound interpretation in the life service domain. The experimental results on our constructed and public datasets prove the effectiveness of our solution. Furthermore, the online A/B testing on Meituan APP suggests that the Query View Click-Through Rate increases by 0.91% when noun compounds are used to enrich semantic information of items with the help of their interpretations on the platform.
Jiaqing Liang, Yanghua Xiao, Fubao Zhang, Zongyu Wang, Rui Xie 0005
IEEE Trans. Knowl. Data Eng.8
2022 An Effective and Efficient Entity Alignment Decoding Algorithm via Third-Order Tensor Isomorphism
abstract
Xin Mao, Meirong Ma, Hao Yuan, Jianchao Zhu, ZongYu Wang, Rui Xie, Wei Wu, Man Lan. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Xin Mao 0002, Meirong Ma, Jianchao Zhu, Zongyu Wang, Man Lan
ACL (1)5
2014 The method for image retrieval based on multi-factors correlation utilizing block truncation coding
Zongyu Wang
Pattern Recognit.2
2013 A novel method for image retrieval based on structure elements' descriptor
Zongyu Wang
J. Vis. Commun. Image Represent.2