VLDB 2026 Research / reviewers in the wild / expert
Ye Wang 0006
dblp:44/6292-6
· DBLP profile ↗
24ranked-venue papers
11as first author
16since 2021 · last 2026
0000-0002-1748-6890ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 6 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 3 since 2021Systems, architecture and hardware · 3Databases, data management, data science and information retrieval · 3 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 1 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-task learning with hierarchical feature disentanglement for low-exposure facial recognition in smart cockpits
Zhuoyi Yu, Guanmeng Xian, Qianmengke Zhao, Hong Yu 0007, Ye Wang 0006 |
Neural Comput. Appl. | 6 |
| 2026 | Visual Evidence-Aware for Object Hallucinations Rectification in LLM-Based Video CaptioningabstractRecent neural models for video captioning are typically built using a framework that combines a pre-trained visual encoder with a large language model(LLM) decoder. However, large language models in video captioning often generate non-existent entities, known as object hallucinations, which severely limit performance. To mitigate object hallucinations, two key issues remain: 1. Biased training data and Knowledge bias in LLM leads models to generate hallucinations; 2. Current methods focus on removal rather than restoring the correct visual content, reducing caption completeness. To address these issues, we propose a visual evidence-aware for object hallucination rectification in LLM-based video captioning. Generally, our model aims to diagnose and correct those generated object hallucinations, and then supplement missing visual content by constraining the process of text description generation. Specifically, we first generate captions by words based on the input video. When decoding each object description, the decoder utilizes visual features for hallucination diagnosis and correction, proposing visual evidence to modify hallucinatory descriptions. This process ensures the generated captions align with the visual content, alleviating the generation of object hallucinations. Compared with the baseline models, our method performs state-of-the-art performance in video captioning, especially avoiding neglecting objects in the visual content caused by the generated hallucinatory descriptions. Ye Wang 0006, Jiancheng Zhou, Qun Liu 0005, Feng Hu 0001, Guoyin Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2026 | Hierarchical Causal Learning for Face Age SynthesisabstractFace age synthesis (FAS) predicts a person's future or past facial appearance. In FAS, modifying one facial attribute usually affects the generation of other attributes during face image generation. Current models directly learn entangled representations of age-related features, resulting in insufficient feature disentanglement, which consequently impairs their causal reasoning capability for FAS tasks. To this end, we propose a hierarchical causal learning model for face age synthesis (HCFace), which integrates hierarchical structures and causal relationships into the facial generative model. Specifically, we propose to leverage hierarchical causal relationships to align with facial features for feature disentanglement. Furthermore, we design a novel nonlinear mapping function that captures the true patterns of facial attribute changes with age, enhancing the disentanglement of these attributes. We conduct extensive experiments to validate the superiority of our proposed model. Compared to other advanced baseline methods, HCFace improves overall accuracy by 2.47%, with improvements of 9.75% and 9.69% in certain age-related attributes, such as skin and hair. Our source code is available at https://github.com/SE-hash/HCFace. Ye Wang 0006, Pan Sun, Lifeng Shen, Jiaxu Leng, Guoyin Wang 0001, Hong Yu 0007 |
IEEE Trans. Image Process. | 1 |
| 2025 | Generating samples for covariance to update prototype in few-shot class-incremental learning
Hong Yu 0007, Qiwei Luo, Ye Wang 0006, Guoyin Wang 0001 |
Appl. Intell. | 3 |
| 2025 | Hierarchical chat-based strategies with MLLMs for Spatio-temporal action detection
Ye Wang 0006, Fei Tao 0003, Hong Yu 0007, Qun Liu 0005 |
Inf. Process. Manag. | 2 |
| 2025 | Dynamic Emotion-Dependent Network With Relational Subgraph Interaction for Multimodal Emotion RecognitionabstractMultimodal Emotion Recognition in Conversations (MERC) is an important topic in human-computer interaction. In the MERC task, conversations exhibit dynamic emotional dependency, including inter-speaker and intra-speaker emotional dependency, both are vital in understanding the content. However, current research primarily integrates these two emotional dependencies into one unified module, limiting the accuracy of MERC. In this paper, we propose a dynamic emotion-dependent network with relational subgraph interaction named DEDNet. DEDNet introduces relational subgraphs to separately model two emotional dependencies, enabling structured learning paths for utterances based on distinct emotional dependency types. Specifically, nodes indicate the utterances at different moments in the conversation, while edges define the emotional dependency and temporal relationships between nodes. To explicitly capture the differences between these two emotional dependencies, distinct subgraphs are designed for comprehensive representations. Furthermore, we propose an incremental interactive strategy, sequentially leveraging two emotional dependencies to learn the changes in dependency relationships. We find that modeling inter-speaker emotional dependency can better identify negative emotions and modeling intra-speaker emotional dependency can better recognize positive emotions. Experimental results demonstrate that our model outperforms current state-of-the-art methods on three benchmark datasets, IEMOCAP, MELD and DailyDialog. Ye Wang 0006, Wei Zhang 0322, Ke Liu 0008, Wei Wu 0022, Feng Hu 0001, Hong Yu 0007, Guoyin Wang 0001 |
IEEE Trans. Affect. Comput. | 1 |
| 2025 | Cross-Interaction of Chinese Characters Structures and Boundary Features for Improving Clinical Named Entity RecognitionabstractIn the natural language processing task of clinical named entity recognition (CNER), accurately identifying the boundaries and categories of medical entities is crucial. However, traditional methods struggle to recognize a large number of clinical terms and symbols that have never been encountered before, ultimately limiting the performance of CNER. Besides, there exist some easy-to-confuse Chinese clinical entities that are semantically similar but belong to quite different categories, such as "" (pulmonary nodules, a symptom entity) and "" (pulmonary tuberculosis, a disease entity), which can lead to entity misidentification. To address these problems, we propose a novel NER model called Cross-Interaction of Chinese characters structures and Boundary Features (CCS). The proposed model leverages Chinese character structural features and boundary information to comprehensively and accurately identify confusing entities. We further design a Cross-Attention mechanism to capture dependency relationships between different entities and radicals of characters, enhancing the model's semantic understanding of specialized terms and symbols, as well as improving its ability to recognize boundaries. Our experimental results show that our proposed model outperforms other state-of-the-art models on various public medical datasets, achieving significant improvements on the CCKS2020, CMeEE, CMI, and IMCS datasets, respectively. Ye Wang 0006, Hong Yu 0007, Guoyin Wang 0001, Chunmeng Shi, Dajiang Lei |
IEEE J. Biomed. Health Informatics | 1 |
| 2024 | Self-supervised modal optimization transformer for image captioning
Ye Wang 0006, Daitianxia Li, Qun Liu 0005, Li Liu 0030, Guoyin Wang 0001 |
Neural Comput. Appl. | 1 |
| 2024 | One-shot knowledge graph completion based on disentangled representation learning
Youmin Zhang 0006, Ye Wang 0006, Qun Liu 0005, Li Liu 0030 |
Neural Comput. Appl. | 3 |
| 2023 | Stable local interpretable model-agnostic explanations based on a variational autoencoder
Xu Xiang, Hong Yu 0007, Ye Wang 0006, Guoyin Wang 0001 |
Appl. Intell. | 3 |
| 2023 | Where to look: Multi-granularity occlusion aware for video person re-identification
Jiaxu Leng, Xinbo Gao 0001, Yan Zhang 0108, Ye Wang 0006, Mengjingcheng Mo |
Neurocomputing | 5 |
| 2023 | Interpreting open-domain dialogue generation by disentangling latent feature representations
Ye Wang 0006, Jingbo Liao, Hong Yu 0007, Guoyin Wang 0001, Li Liu 0030 |
Neural Comput. Appl. | 1 |
| 2022 | RCNet: Recurrent Collaboration Network Guided by Facial Priors for Face Super-ResolutionabstractIn this paper, we present a novel FSR method, called RCNet, which progressively improves the performance of FSR and Landmark Estimation (LE) via recurrent collaboration. In our approach, FSR and LE complement each other. Different from previous FSR methods that directly estimate the facial landmarks on the low-resolution face images, the proposed RCNet conducts LE on the super-resolution face image obtained through multiple iterations. Benefiting from the super-resolution face images, facial landmarks are precisely estimated, which boosts the performance of FSR in turn. Furthermore, we design a Component-Aware Fusion Module (CAFM) for better recovering facial details, which adaptively fine-tunes the estimated landmarks and groups them into face components to maximize the guiding role of the facial land-marks. In addition, the iterative feature aggregation is developed to preferably capture the information from the LR/SR face images. Experimental results show that the proposed RCNet outperforms the state-of-the-art methods in both quantitative and qualitative aspects for super-resolving very low-resolution faces. Jiaxu Leng, Ye Wang 0006 |
ICME | 2 |
| 2022 | LRP2A: Layer-wise Relevance Propagation based Adversarial attacking for Graph Neural Networks
Li Liu 0030, Ye Wang 0006, William Kwok-Wai Cheung, Youmin Zhang 0006, Qun Liu 0005, Guoyin Wang 0001 |
Knowl. Based Syst. | 3 |
| 2022 | Semantic-aware conditional variational autoencoder for one-to-many dialogue generation
Ye Wang 0006, Jingbo Liao, Hong Yu 0007, Jiaxu Leng |
Neural Comput. Appl. | 1 |
| 2021 | Realize your surroundings: Exploiting context information for small object detection
Jiaxu Leng, Yihui Ren 0002, Xiaoding Sun, Ye Wang 0006 |
Neurocomputing | 5 |
| 2020 | Attention augmentation with multi-residual in bidirectional LSTM
Ye Wang 0006, Xinxiang Zhang, Mi Lu, Yoonsuck Choe |
Neurocomputing | 1 |
| 2019 | An Attention-aware Bidirectional Multi-residual Recurrent Neural Network (Abmrnn): A Study about Better Short-term Text ClassificationabstractLong Short-Term Memory (LSTM) has been proven an efficient way to model sequential data, because of its ability to overcome the gradient diminishing problem during training. However, due to the limited memory capacity in LSTM cells, LSTM is weak in capturing long-time dependency in sequential data. To address this challenge, we propose an Attention-aware Bidirectional Multi-residual Recurrent Neural Network (ABMRNN) to overcome the deficiency. Our model considers both past and future information at every time step with omniscient attention based on LSTM. In addition to that, the multi-residual mechanism has been leveraged in our model which aims to model the relationship between current time step with further distant time steps instead of a just previous time step. The results of experiments show that our model achieves state-of-the-art performance in classification tasks. Ye Wang 0006, Xinxiang Zhang, Theodora Chaspari, Yoonsuck Choe, Mi Lu |
ICASSP | 1 |
| 2019 | A New Method for Sparse Signals Reconstruction in Fusion Center with Incomplete MeasurementsabstractIn this paper, we focus on a specific problem in joint sparse signals recovery, where not all of the measurements are fully received (observed) in the data fusion center (FC), but some of them are missing. Most of the previous work focused on the sparse signals recovery problem that the information of the measurements in FC is complete (fully observed). However, in practice, usually it is not easy for the data center to receive complete measurements, due to limited sampling rate, noise corruption, or transmission error, etc. Hence, how to reconstruct the sparse signals under the incomplete measurements is an important problem for the fusion center. Based on Regularized M-FOCUSS (FOCal Underdetermined System Solver) algorithm, this paper proposes a new method to efficiently reconstruct the sparse signals from the incomplete measurements. Different to the widely used matrix completion based method, our method obtains the solution from the received measurements directly, instead of running the matrix completion algorithm in advance. Furthermore, our method shows excellent recovery performance compared to the matrix completion based method, and also the independent vector recovery method, and the block-sparse based recovery method. Experimental results are also provided to show the effectiveness of our proposed scheme. Jiafan Wang 0002, Ye Wang 0006, Zhengguang Zheng |
ICC | 3 |
| 2019 | English Out-of-Vocabulary Lexical Evaluation TaskabstractUnlike previous unknown nouns tagging task, this is the first attempt to focus on out-of-vocabulary (OOV) lexical evaluation tasks that does not require any prior knowledge. The OOV words are words that only appear in test samples. The goal of tasks is to provide solutions for OOV lexical classification and predication. The tasks require annotators to conclude the attributes of the OOV words based on their related contexts. Then, we utilize unsupervised word embedding methods such as Word2Vec and Word2GM to perform the baseline experiments on the categorical classification task and OOV words attribute prediction tasks. Ye Wang 0006, Xinxiang Zhang, Mi Lu, Yoonsuck Choe, Jingjing Cao |
INDIN | 2 |
| 2019 | Bend Detection of Bridge Chords in UAV Images via Region-Based Deep Semantic Segmentation NetworkabstractVision-based techniques are gradually being applied to civil engineering fields by aiding human resources to inspect and assess the condition of structures. This paper presents a vision-based inspection system to detect bends of the bridge chords in UAV images. The main parts of this paper firstly present a novel region proposal algorithm to localize bridge chord bends without distinct boundaries and then present a fusion-based deep semantic segmentation network that improves the detection performance by fusing spatial information from neighborhood. The proposed system achieves approximate 0.7762 Mean IOU and 0.6424 BF Score on the field dataset, which outperforms the other state-of-the-art systems. Xinxiang Zhang, Ye Wang 0006, Yue Zhang 0004, Hao Wu 0059, Ming-Bo Zhao |
INDIN | 2 |
| 2018 | A deep learning model for automatic evaluation of academic engagementabstractThis paper proposed a deep learning model for automatic evaluation of academic engagement based on video data analysis. A coding system based on the BROMP standard for behavioral, emotional, and cognitive states was defined to code typical videos in an autonomous learning environment. Then after the key points of human skeletons were extracted from these videos using pose estimation technology, deep learning methods were used to realize the effective recognition and judgment of motion and emotions. Based on this, an analysis and evaluation of learners' learning states was accomplished, and a prototype of academic engagement evaluation system was successfully established eventually. Ye Wang 0006, Weining Qian, Aoying Zhou |
L@S | 3 |
| 2009 | Knowledge Discovery from Academic Search Engine
Ye Wang 0006, Xiaoling Wang 0004, Aoying Zhou |
KSEM | 1 |
| 2007 | Academic web search engine: generating a survey automaticallyabstractGiven a document repository, search engine is very helpful to retrieve information. Currently, vertical search is a hot topic, and Google Scholar [4] is an example for academic search. However, most vertical search engines only return the flat ranked list without an efficient result exhibition for given users. We study this problem and designed a vertical search engine prototype Dolphin, where the flexible user-oriented templates can be defined and the survey-like results are presented according to the template. Ye Wang 0006, Zhihua Geng, Xiaoling Wang 0004, Aoying Zhou |
WWW | 1 |