VLDB 2026 Research / reviewers in the wild / expert
Yating Yang
dblp:47/11349
· DBLP profile ↗
50ranked-venue papers
5as first author
34since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 25 · 13 since 2021Computer networks · 12 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 7 since 2021Databases, data management, data science and information retrieval · 5 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Systems, architecture and hardware · 1Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | M3UCD: A Multi-task Multimodal Metaphor Understanding Challenge Dataset for LLMsabstractUnderstanding multimodal metaphors represents a crucial pathway for machines to comprehend human cognition. However, current research remains constrained by superficial dataset annotations, insufficient systematic evaluation of large language models, and fragmented task frameworks. To bridge these gaps, the paper proposes a systematic solution featuring: (I) We present the largest fine-grained Multi-task Multimodal Metaphor Understanding Challenge Dataset (M3UCD) built via multi-perspective collaborative annotation. It contains 15,345 samples, each annotated with 12 manual attribute labels. (II) Systematic benchmarking of LLMs' capacity boundaries in metaphor understanding. Evaluation results reveal the persistent challenges LLMs face in this domain while validating M3UCD's effectiveness and potential. (III) A concise and unified multi-task baseline framework was developed and demonstrated its effectiveness in enhancing the metaphor understanding capabilities of MLLMs. Tianlong Zheng, Yating Yang, Rui Dong 0002, Bo Ma 0004, Lei Wang 0065, Xi Zhou 0007, Siru Miao, Osman Turghun |
AAAI | 2 |
| 2026 | Flickr-ABS: A Benchmark for Abstract Intent in Image-Text RetrievalabstractIn image–text retrieval, users often express abstract, incomplete, or polysemous intents rather than fully concrete visual descriptions. In contrast, existing retrieval benchmarks predominantly rely on object- and action-centric captions that presuppose detailed and visually grounded descriptions, leaving model behavior under semantically fuzzy queries largely underexplored. To address this gap, we introduce Flickr-ABS, an abstraction-aware image–text retrieval benchmark built upon Flickr30k. For each image, Flickr-ABS associates retrieval queries at three explicitly defined semantic levels, ranging from concrete scene descriptions to high-level conceptual interpretations. Extensive evaluations on a wide range of vision–language models reveal a substantial degradation in retrieval performance as query abstraction increases, highlighting the fragility of current retrieval systems under abstract semantic intent. We further demonstrate that jointly leveraging supervision from multiple abstraction levels can improve retrieval robustness for abstract queries while preserving performance on concrete ones. We further propose a semantic recall protocol that accounts for conceptual equivalence, enabling more faithful evaluation under high-level abstraction. Together, Flickr-ABS and the proposed evaluation protocol offer a practical benchmark for analyzing and improving abstraction-aware image–text retrieval. Our code is available at https://github.com/LT1anY/Flickr-abs. Tianyuan Li, Lei Wang 0065, Yating Yang, Rui Dong 0002, Ahtamjan Ahmat, Bangju Han |
ICMR | 3 |
| 2026 | Ask-to-Retrieve: VQA-Guided Information-Gain Question Selection for Interactive Text-to-Image RetrievalabstractInteractive image retrieval iteratively interacts with users to capture retrieval intent and refine textual queries, effectively mitigating the ambiguity and performance degradation inherent in single-round text-to-image retrieval. However, existing LLM-based methods largely rely on image captions or dialogue history to extract features within a limited candidate space. Consequently, they often generate questions that are either irrelevant to the target’s distinctive characteristics or fail to capture fine-grained visual differences. Furthermore, an effective question should efficiently distinguish the target from non-target samples—a critical aspect overlooked by prior works, resulting in suboptimal retrieval efficiency. To bridge this gap, we present A2R (Ask-to-Retrieve), a visually grounded framework designed to maximize information gain. Specifically, a Diversity Candidate Perception (DCP) module constructs a compact visual grid to capture diverse visual semantics. Subsequently, a Discriminative Question Generation (DQG) module observes this grid to propose discriminative questions. Finally, a Maximum Information Gain Decision (IGD) module evaluates these questions and selects the one maximizing entropy reduction. Experiments on VisDial, COCO, and Flickr30k demonstrate that our method achieves superior retrieval accuracy and efficiency compared to strong baselines. Shuaiwei Xie, Yating Yang, Bo Ma 0004, Xi Zhou 0007, Zhen Wang 0062, Ahtamjan Ahmat |
ICMR | 2 |
| 2026 | IC-Lite: Enabling Lightweight and Real-Time Data Dissemination Over Information-Centric IoTabstractInformation-Centric Internet of Things (IC-IoT) enhances content distribution through its inherent in-network caching and multicast capabilities. However, effective data naming and disseminating in dynamic IoT environments remain challenging. In particular, each IoT device requires a routable name to facilitate data production and retrieval. Assigning and maintaining routable names for large numbers of sensors is cumbersome, restricting data management flexibility and scalability. Moreover, while data pushing is essential for IoT applications, such as real-time data uploads, the receiver-driven nature of ICN hinders push-based communication. In this context, we propose IC-Lite, an optimized framework built on IC-IoT, designed to provide lightweight, information-centric communication with simplified name assignment and data provisioning. Our approach introduces a modified forwarding plane that transmits data using simplified names and pushes data with the three-packet handshake, thereby enabling timely and efficient dissemination of real-time data from devices to application proxies. Evaluation results and theoretical analysis demonstrate that IC-Lite significantly reduces the overhead associated with name-based routing and namespace management, achieving an order of magnitude improvement in efficiency. Additionally, IC-Lite reduces data delivery delay by 60.8% in dynamic scenarios, making it an efficient solution for real-time IoT applications. Yating Yang, Jinglan Song, Tian Song 0004 |
IEEE Internet Things J. | 1 |
| 2025 | Missed Abortion Prediction Using Dual Network on Large-Scale Clinical Data with Multiple FeaturesabstractMissed abortion is a common obstetric complication with serious consequences for maternal health and fertility. Accurately predicting the risk of missed abortion occurrence is important for timely intervention and prevention. Given the large amount and complexity of clinical data, as well as the nonlinear inter-correlations that exist among multiple features, it is difficult for general machine learning methods to effectively mine the potential crucial patterns that dominate the occurrence of missed abortions. To address this problem, this paper proposes a dual network framework, namely DNMAP, which integrates a local feature extraction module to capture correlations among clinical indicators with a feature integration module to aggregate these features for missed abortion prediction. In addition, DNMAP extends the prediction problem to a multi-classification task to better explore the association between missed abortion and recurrent miscarriage where two or more miscarriages have occurred. Different from previous studies, this paper is the first to investigate the problem of missed abortion in this way, with the largest practical clinical miscarriage dataset to our knowledge. A series of experiments have been conducted and the results demonstrate that the proposed DNMAP significantly outperforms various types of models in the task of missed abortion prediction. DNMAP can not only predict the risk of missed abortion a simple and effective manner, but also provide more refined risk stratification, which suggests novel insights and methods to further improve maternal health and reproductive outcomes. Jun Zhang 0003, Xiaoli Bo, Yue Yang 0035, Xiwen Yang, Yating Yang |
BIBM | 7 |
| 2025 | Low-Resource Language Expansion and Translation Capacity Enhancement for LLM: A Study on the UyghurabstractAlthough large language models have significantly advanced natural language generation, their potential in low-resource machine translation has not yet been fully explored, especially for languages that translation models have not been trained on. In this study, we provide a detailed demonstration of how to efficiently expand low-resource languages for large language models and significantly enhance the model’s translation ability, using Uyghur as an example. The process involves four stages: collecting and pre-processing monolingual data, conducting continuous pre-training with extensive monolingual data, fine-tuning with less parallel corpora using translation supervision, and proposing a direct preference optimization based on translation self-evolution (DPOSE) on this basis. Extensive experiments have shown that our strategy effectively expands the low-resource languages supported by large language models and significantly enhances the model’s translation ability in Uyghur with less parallel data. Our research provides detailed insights for expanding other low-resource languages into large language models. Kaiwen Lu, Yating Yang, Fengyi Yang, Rui Dong 0002, Bo Ma 0004, Aihetamujiang Aihemaiti, Abibulla Atawulla, Lei Wang 0065, Xi Zhou 0007 |
COLING | 2 |
| 2025 | OpenForecast: A Large-Scale Open-Ended Event Forecasting DatasetabstractComplex events generally exhibit unforeseen, multifaceted, and multi-step developments, and cannot be well handled by existing closed-ended event forecasting methods, which are constrained by a limited answer space. In order to accelerate the research on complex event forecasting, we introduce OpenForecast, a large-scale open-ended dataset with two features: (1) OpenForecast defines three open-ended event forecasting tasks, enabling unforeseen, multifaceted, and multi-step forecasting. (2) OpenForecast collects and annotates a large-scale dataset from Wikipedia and news, including 43,419 complex events spanning from 1950 to 2024. Particularly, this annotation can be completed automatically without any manual annotation cost. Meanwhile, we introduce an automatic LLM-based Retrieval-Augmented Evaluation method (LRAE) for complex events, enabling OpenForecast to evaluate the ability of complex event forecasting of large language models. Finally, we conduct comprehensive human evaluations to verify the quality and challenges of OpenForecast, and the consistency between LEAE metric and human evaluation. OpenForecast and related codes will be publicly released. Zhen Wang 0062, Xi Zhou 0007, Yating Yang, Bo Ma 0004, Lei Wang 0065, Rui Dong 0002, Azmat Anwar |
COLING | 3 |
| 2025 | Mining the Past with Dual Criteria: Integrating Three types of Historical Information for Context-aware Event ForecastingabstractEvent forecasting requires modeling historical event data to predict future events, and achieving accurate predictions depends on effectively capturing the relevant historical information that aids forecasting.Most existing methods focus on entities and structural dependencies to capture historical clues but often overlook implicitly relevant information.This limitation arises from overlooking event semantics and deeper factual associations that are not explicitly connected in the graph structure but are nonetheless critical for accurate forecasting.To address this, we propose a dual-criteria constraint strategy that leverages event semantics for relevance modeling and incorporates a self-supervised semantic filter based on factual event associations to capture implicitly relevant historical information.Building on this strategy, our method, termed ITHI (Integrating Three types of Historical Information), combines sequential event information, periodically repeated event information, and relevant historical information to achieve context-aware event forecasting.We evaluated the proposed ITHI method on three public benchmark datasets, achieving state-of-the-art performance and significantly outperforming existing approaches.Additionally, we validated its effectiveness on two structured temporal knowledge graph forecasting dataset 1 . Lei Wang 0065, Yating Yang, Bo Ma 0004, Rui Dong 0002, Fengyi Yang, Ahtamjan Ahmat, Kaiwen Lu |
EMNLP | 3 |
| 2025 | Demand-Oriented Robust Service Allocation with Multi-Phase Task Matching in Satellite NetworksabstractThe rapid growth in satellite internet services, exemplified by constellations like Starlink, has underscored the imperative for efficient service allocation in satellite networks. This poses significant challenges due to the dynamic and multidimensional nature of user demands, which vary across time and space, alongside the fluctuating availability of satellite resources. Traditional static allocation methods are inadequate for these dynamic conditions, often resulting in inefficient resource utilization and compromised service quality. To address these challenges, this paper presents Demand-Oriented Robust Service Allocation (DORSA) with a multi-phase task matching approach. Our method begins with a static matching phase to establish an initial user demand-service match, followed by a dynamic matching phase that leverages a sparse time-varying bipartite graph model to enhance computational efficiency. The DORSA algorithm incorporates local backup-aware matching to rapidly adapt to demand fluctuations and satellite movements, ensuring continuous and high-quality service provision. Simulation results indicate that our approach reduces the average response latency by 74.5% compared to the baseline approach, and by 54.2% compared to the state-of-the-art task allocation algorithm under dynamic demand conditions within satellite networks. Yating Yang, Huanyu Sang |
ICC | 2 |
| 2025 | LiteCobra: Enhancing Java Deserialization Vulnerability Detection with Call Graph PruningabstractJava deserialization vulnerabilities have become a critical security threat, challenging to detect and even harder to exploit due to deserialization's flexible and customizable nature. Researchers have proposed code property graph methods for comprehensive program analysis and controllability analysis algorithms to filter out irrelevant method calls. However, existing approaches often face limitations in accuracy and efficiency, especially in complex Java deserialization scenarios, leaving a gap in fully addressing this security issue. To address the challenges in detecting Java deserialization gadget chains, this paper proposes LiteCobra—a method that employs edge-cut optimization for rapid call graph construction and uses a controllability analysis algorithm for efficient pruning during Java deserialization gadget chain discovery. Our tests on the ysoserial dataset show that LiteCobra reduces the false positive rate by 67.5 % and enhances efficiency by 49.6 % compared to state-of-the-art approaches for detecting Java deserialization vulnerabilities. These results indicate that LiteCobra significantly enhances efficiency while maintaining high accuracy in detecting Java deserialization gadget chains. Yating Yang |
ICC | 2 |
| 2025 | Bidirectional Feature Fusion and Adaptive Decision Network for Multimodal Fake News DetectionabstractThe proliferation of multimedia content on social media platforms has rendered the detection of fake news an increasingly critical challenge. Despite the progress made in multimodal fake news detection methods, significant challenges remain: information loss during the cross-modal feature fusion process and restricted analysis due to single-perspective decision mechanisms. We propose the Bidirectional Feature Fusion and Adaptive Decision Network (BFF-ADN), which introduces two key innovations: a bidirectional token-level feature fusion mechanism that captures key correlations between visual and textual modalities while reducing information loss, and an Adaptive Multi-perspective Decision Integration Network (AMDI-Net) that integrates multiple specialized detectors for comprehensive analysis. Experiments on three real-world datasets (Weibo, GossipCop, and PolitiFact) demonstrate that BFF-ADN achieves state-of-the-art performance. Ablation studies and visualizations confirm the effectiveness of our proposed components, particularly highlighting the advantages of the bidirectional feature fusion mechanism and Adaptive Multi-perspective Decision Integration Network. Dilxat Abdureyim, Bo Ma 0004, Yating Yang, Rui Dong 0002, YiDu Chen, Azmat Anwar, Lei Wang 0065 |
ICME | 3 |
| 2025 | LyRE: Learning Varying Fusion Degrees with Hierarchical Aggregation to Improve Multimodal Misinformation Detection
YiDu Chen, Bo Ma 0004, Yating Yang, Dilxat Abdureyim, Rui Dong 0002, Zhen Wang 0062, Lei Wang 0065, Xi Zhou 0007 |
NLPCC (2) | 3 |
| 2025 | Self-Distillation Across Modalities: Enhancing Cross-Modal Correlation Perception for Multimodal Fake News Detection
YiDu Chen, Bo Ma 0004, Yating Yang, Dilxat Abdureyim, Chirui Zhang, Rui Dong 0002, Lei Wang 0065, Xi Zhou 0007 |
NLPCC (3) | 3 |
| 2025 | ISIPN: Intention-Semantic Incongruity Perception Network for Multimodal Metaphor Detection
Tianlong Zheng, Yating Yang, Rui Dong 0002, Bo Ma 0004, Lei Wang 0065, Xi Zhou 0007 |
NLPCC (3) | 2 |
| 2025 | PtCAM: Scalable High-Speed Name Prefix Lookup using TCAMabstractName prefix lookup is a core characteristic of name-based packet forwarding in ICN and also URL redirections in CDN. This operation is essentially a longest prefix match on URL-like names, which faces significant challenges in high-speed processing due to large set, length-unbounded names and per-packet lookup need. In this paper, we propose a scheme to support generic hierarchical name lookups on millions of prefixes with fast, constant and guaranteed performance using ternary content-addressable memory (TCAM). Our work has three contributions. First, a scheme, namely PtCAM, is proposed to exploit TCAM for name lookup using a binary Patricia trie. Second, several approaches are proposed to enhance efficiency and scalability. Third, experiments on various name datasets are carried out. The results show that our scheme has the potential to achieve high-speed lookup performance nearly equal to the given TCAM throughput. Notably, PtCAM can be implemented on TCAM-based IP linecards, which means that industrial commercial off-the-shelf devices can be directly involved in the advancement of information-centric innovation and deployment. Tianlong Li, Yating Yang |
SIGCOMM | 3 |
| 2025 | FOFL: Dynamic Function Output-Based Software Fault Localization via Deep LearningabstractLearning-based fault localization has become a prominent research direction in software engineering due to its ability to leverage diverse program artifacts such as execution traces and coverage data for precise fault identification. However, current techniques face two fundamental limitations: (1) oversimplified representations of program behavior through basic coverage metrics, and (2) high computational cost associated with collecting fine-grained runtime data. In this work, we present FOFL, a novel function output-based fault localization approach. Our intuition is that programs can be modeled as complex dynamical systems, where faults manifest as perturbations observable through function outputs. Our method first captures function-level output patterns during execution, then encodes them into a structured matrix representation that preserves system-level behavioral signatures. These matrices are analyzed through a hybrid deep learning architecture combining convolutional neural networks for spatial feature extraction and attention mechanisms for critical pattern recognition, ultimately producing a ranked list of suspicious locations. Experiments on the widely used Defects4J benchmark show that FOFL outperforms existing techniques by localizing 204 bugs in the Top-1 ranking, which corresponds to a 25-bug improvement over state-of-the-art methods, while also enhancing MFR and MAR by 4.5% and 4.3%, respectively. Furthermore, ablation studies confirm the positive contribution of listwise loss function and feature aggregation design. The results establish FOFL as an effective solution that advances fault localization through principled system modeling and targeted learning of informative data representations. Qingyuan Sun, Tu Peng, Yating Yang |
SMC | 3 |
| 2025 | M2-VLP: Enhancing Multilingual Vision-Language Pre-Training via Multi-Grained AlignmentabstractRecently, multilingual Vision-Language Pre-training (mVLP) has shown remarkable progress in learning joint representations across different modalities and languages. However, most existing methods learn semantic alignment at a coarse-grained level and fail to capture fine-grained correlations between different languages and modalities. To address this, we propose a Multi-grained Multilingual Vision-Language Pre-training (M2-VLP) model, which aims to learn cross-lingual cross-modal alignment at different semantic granular levels. In cross-lingual interaction, the model learns the global alignment of parallel sentence pairs and the word-level correlations. In cross-modal interaction, the model aligns images with captions and image regions with corresponding words. To integrate the cross-lingual and cross-modal alignment above, we propose a unified multi-grained contrastive learning paradigm. Under zero-shot cross-lingual and fine-tuned multilingual settings, extensive experiments on vision-language downstream tasks across twenty languages demonstrate the effectiveness of M2-VLP over competitive contrastive models. Code and models are available at https://github.com/ahtamjan/M2-VLP. Ahtamjan Ahmat, Lei Wang 0065, Yating Yang, Bo Ma 0004, Rui Dong 0002, Kaiwen Lu |
WWW | 3 |
| 2025 | Cross-lingual prompting method with semantic-based answer space clustering
Ahtamjan Ahmat, Yating Yang, Bo Ma 0004, Rui Dong 0002, Lei Wang 0065 |
Appl. Intell. | 2 |
| 2025 | SGEU: enhancing LLM reasoning via backward exemplar generation and verification
Zhen Wang 0062, Xi Zhou 0007, Yating Yang, Bo Ma 0004, Lei Wang 0065, Rui Dong 0002 |
Appl. Intell. | 3 |
| 2025 | An Empirical Study of MBFL on Novice Programs Across Different Programming LanguagesabstractProgramming education in computer science is growing rapidly, and debugging is a key challenge for novice programmers due to their limited experience. Mutation-Based Fault Localization (MBFL) is widely used in industry, but its effectiveness and challenges in novice programs need further study. While Python is a popular language in machine learning and data science, there is little research comparing fault localization in Python and Java for novice programmers. To bridge this gap, we conduct an empirical study to evaluate MBFL’s accuracy and execution overhead in common novice programming errors across different languages. We analyze how program features like code coverage and mutation score affect MBFL’s performance and whether these effects differ between languages. We also examine how MBFL’s effectiveness changes when suspiciousness scores are the same and how mutant noise and coincidental correct test cases vary across languages. Additionally, we propose a mutation confidence formula based on repair potential and behavioral difference to assess the usefulness of mutants in MBFL. Our study demonstrates that MBFL works well for novice fault localization in both Java and Python, with Python performing better. MBFL correctly identifies 45, 70, and 92 faults within the TOP-N (N = 1, 3, 5), proving its strong performance. However, tie problems, mutant noise, and coincidental correct test cases weaken MBFL, especially in Java. Results in both languages show a strong positive correlation between mutant confidence and fault localization accuracy, confirming the formula’s effectiveness across languages. Yating Yang, Ao Mei, Chao Yang 0043 |
Int. J. Softw. Eng. Knowl. Eng. | 1 |
| 2024 | Pruning Residual Networks in Multilingual Neural Machine Translation to Improve Zero-Shot Translation
Kaiwen Lu, Yating Yang, Rui Dong 0002, Bo Ma 0004, Lei Wang 0065, Xi Zhou 0007, Ahtamjan Ahmat |
NLPCC (3) | 2 |
| 2024 | iStack: A General and Stateful Name-based Protocol Stack for Named Data Networking
Tianlong Li, Yating Yang |
NSDI | 3 |
| 2024 | TokenFree: A Tokenization-Free Generative Linguistic Steganographic Approach with Enhanced ImperceptibilityabstractSince tokenization serves a fundamental preprocessing step in numerous language models, tokens naturally constitute the basic embedding units for generative linguistic steganography. However, tokenization-based methods face challenges including limited embedding capacity and possible segmentation ambiguity. Despite existing character-level (one tokenization-free type) linguistic steganographic approaches, they face the problem of generating unknown or out-of-vocabulary words, potentially compromising steganographic imperceptibility. In this paper, we focus on both embedding capacity and imperceptibility of tokenization-free linguistic steganography. First, we suggest that unknown words mainly result from low-entropy distributions and rigid coding rules used in candidate pools, thus we propose an entropy-based selection approach to flexibly construct candidate pools. Further, we present a lexical emphasis approach, prioritizing characters within candidate pools capable of forming in-vocabulary words. Experiments show that, across a range of high embedding rates, our approaches achieve considerably higher imperceptibility and text fluency, increase anti-steganalysis capacity averagely by 14.4%, and particularly reduce out-of-vocabulary rate averagely by 88.7%, compared to the existing state-of-the-art character-level steganographic methods. Ruiyi Yan, Yating Yang |
SMC | 3 |
| 2024 | A Near-Imperceptible Disambiguating Approach via Verification for Generative Linguistic SteganographyabstractGenerative linguistic steganography aims to embed information into natural language texts to achieve covert transmission. However, currently in most approaches based on subword-supporting language models, the extraction process relies on tokenizing steganographic texts into tokens, which could cause segmentation ambiguity, leading to false results or failures of extraction finally. Despite several existing countermeasures (or disambiguation) that have been proposed, they are based on removing tokens of candidate pools, which render them incompatible from the sights of keeping imperceptibility, potentially incurring safety risks. To avoid it, we focus on tackling segmentation ambiguity with near-integrity of candidate pools. In this paper, we propose a near-imperceptible disambiguating approach via verification for generative linguistic steganography. First, this paper draws an all-case extraction method to obtain possible true extracted results. Further, length verification and checksum verification are presented to filter wrong extracted results caused by segmentation ambiguity. Experiments show that our disam-biguating approach outperforms the existing disambiguating approaches, on various criteria, including about 23.49 % higher embedding capacity, about 23.46 % higher imperceptibility and about 5.73% anti-steganalysis capacity of steganographic texts. Ruiyi Yan, Yating Yang |
SMC | 3 |
| 2024 | Relational concept enhanced prototypical network for incremental few-shot relation classification
Bo Ma 0004, Lei Wang 0065, Xi Zhou 0007, Zhen Wang 0062, Yating Yang |
Knowl. Based Syst. | 6 |
| 2024 | Adaptive Segmented Subscription for Efficient Data Dissemination in Vehicular Named Data NetworksabstractVehicular Named Data Networking (VNDN) is a promising information-centric network architecture to achieve efficient data delivery among vehicles. With a subscription-based communication paradigm, it can disseminate event-triggered data like traffic updates or road accident notifications in a timely way. However, due to vehicle mobility and network dynamics, the data subscription path is normally fragile and consequently the data dissemination efficiency would be diminished. To address this problem, we propose SegSub, an adaptive segmented subscription mechanism which can maintain robust subscription path for efficient data dissemination in mobile scenarios. SegSub divides the whole subscription path into several segments and different strategies are employed to maintain subscription status for each segment path. The frequency of subscription update on each segment is calculated based on the prediction of vehicle mobility and network status to reach a trade-off between subscription updating overhead and expected data dissemination delay. Furthermore, a subscription migration mechanism is proposed to alleviate the data loss and redundant pending overhead caused by subscription failure when vehicles move. Evaluation results demonstrate that SegSub can decrease dissemination delay by 91.9% and 58.7% compared to native VNDN and the state-of-the-art subscription method. Jinglan Song, Yating Yang, Wenyi Jin, Tian Song 0004 |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2023 | A Domain-Transfer Meta Task Design Paradigm for Few-Shot Slot TaggingabstractFew-shot slot tagging is an important task in dialogue systems and attracts much attention of researchers. Most previous few-shot slot tagging methods utilize meta-learning procedure for training and strive to construct a large number of different meta tasks to simulate the testing situation of insufficient data. However, there is a widespread phenomenon of overlap slot between two domains in slot tagging. Traditional meta tasks ignore this special phenomenon and cannot simulate such realistic few-shot slot tagging scenarios. It violates the basic principle of meta-learning which the meta task is consistent with the real testing task, leading to historical information forgetting problem. In this paper, we introduce a novel domain-transfer meta task design paradigm to tackle this problem. We distribute a basic domain to each target domain based on the coincidence degree of slot labels between these two domains. Unlike classic meta tasks which only rely on small samples of target domain, our meta tasks aim to correctly infer the class of target domain query samples based on both abundant data in basic domain and scarce data in target domain. To accomplish our meta task, we propose a Task Adaptation Network to effectively transfer the historical information from the basic domain to the target domain. We carry out sufficient experiments on the benchmark slot tagging dataset SNIPS and the name entity recognition dataset NER. Results demonstrate that our proposed model outperforms previous methods and achieves the state-of-the-art performance. Fengyi Yang, Xi Zhou 0007, Yating Yang, Bo Ma 0004, Rui Dong 0002, Abibulla Atawulla |
AAAI | 3 |
| 2023 | Content-Aware Proportional Caching for Efficient Data Delivery over Satellite NetworkabstractThe Low Earth Orbit (LEO) satellite network has emerged as a crucial infrastructure for global content delivery service, due to its global coverage and low latency. However, the high mobility of satellites and the limited on-board resources impede its data distribution efficiency. In this paper, we propose CA-prop, a content-aware proportional caching scheme to improve the efficiency of content delivery over limited on-board cache resources. CA-prop captures the time-varying requesting feature of contents and leverage this metric to determine the caching proportion of a specific type of content. Based on the satellite weighted time-varying graph, a cooperative cache placement scheme is developed to cache the proportional content along the data delivery path. Furthermore, with prediction of satellite mobility and data delivery direction, we enhance CA-prop with a directional pre-caching strategy to increase the caching service residence time on the mobile large-scale satellite constellation. Our experiment results demonstrate that our strategy is effective in enhancing content distribution efficiency, with an average reduction of 43% in data delivery delay and 90% improvement in cache hit ratio under various content scenarios as compared to state-of-the-art approaches. Jiaran Zhang, Yating Yang, Huanyu Sang, Zhuoqun Gao |
GLOBECOM | 2 |
| 2023 | A Slot-Shared Span Prediction-Based Neural Network for Multi-Domain Dialogue State TrackingabstractThere are a large number of candidate values shared among slots in multi-domain dialogue state tracking (DST). The existing span prediction-based DST methods generally adopt slot-independent value extraction architecture, which ignore the value sharing. Besides, the slot-independent design leads to poor scalability. In this paper, we propose a Slot-shared Span Prediction based Network (SSNet) with a general value extraction module for all slots to tackle these problems. To ensure that the value extraction module is able to distinguish different slots, we introduce a Dynamic Fusion Mechanism (DFM) to extract different slot-aware features. DFM plays the routing role, highlighting different dialogue context tokens for different slots. Specifically, DFM firstly calculates similarity matrixes between the dialogue context and different slots, and then determines important dialogue context token with respect to each slot. Experimental results demonstrate that SSNet outperforms the existing start-of-the-art models on both MultiWOZ 2.1 and MultiWOZ 2.2 datasets. Abibulla Atawulla, Xi Zhou 0007, Yating Yang, Bo Ma 0004, Fengyi Yang |
ICASSP | 3 |
| 2023 | A Secure and Disambiguating Approach for Generative Linguistic SteganographyabstractSegmentation ambiguity in generative linguistic steganography could induce decoding errors. One existing dis-ambiguating way is removing the tokens whose mapping words are the prefixes of others in each candidate pool. However, it neglects probability distribution of candidates and degrades imperceptibility. To enhance steganographic security, meanwhile addressing segmentation ambiguity, we propose a secure and disambiguating approach for linguistic steganography. In this letter, we focus on two questions: (1) Which candidate pools should be modified? (2) Which tokens should be retained? Firstly, we propose a secure token-selection principle that the sum of selected tokens' probabilities is positively correlated to statisti-cal imperceptibility. To meet both disambiguation and optimal security, we present a lightweight disambiguating approach that is finding out a maximum weight independent set (MWIS) in one candidate graph only when candidate-level ambiguity occurs. Experiments show that our approach outperforms the existing method in various security metrics, improving 25.7% statistical imperceptibility and 11.2% anti-steganalysis capacity averagely. Ruiyi Yan, Yating Yang, Tian Song 0004 |
IEEE Signal Process. Lett. | 2 |
| 2023 | WAD-X: Improving Zero-shot Cross-lingual Transfer via Adapter-based Word AlignmentabstractMultilingual pre-trained language models (mPLMs) have achieved remarkable performance on zero-shot cross-lingual transfer learning. However, most mPLMs implicitly encourage cross-lingual alignment in pre-training stage, making it hard to capture accurate word alignment across languages. In this paper, we propose Word-align ADapters for Cross-lingual transfer (WAD-X) to explicitly align word representations of mPLMs via language-specific subspace. Taking a mPLM as the backbone model, WAD-X constructs subspace for each source-target language pair via adapters. The adapters use statistical alignment as the prior knowledge to guide word-level aligning in the corresponding bilingual semantic subspace. We evaluate our model across a set of target languages on three zero-shot cross-lingual transfer tasks: part-of-speech tagging (POS), dependency parsing (DP), and sentiment analysis (SA). Experimental results demonstrate that our proposed model improves zero-shot cross-lingual transfer on three benchmarks, with improvements of 2.19, 2.50, and 1.61 points in POS, DP, and SA tasks over strong baselines. Ahtamjan Ahmat, Yating Yang, Bo Ma 0004, Rui Dong 0002, Kaiwen Lu, Lei Wang 0065 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2023 | Generic and Sensitive Anomaly Detection of Network Covert Timing ChannelsabstractNetwork covert timing channels can be maliciously used to exfiltrate secrets, coordinate attacks and propagate malwares, posing serious threats to cybersecurity. Current covert timing channels normally conduct small-volume transmission under the covers of various disguising techniques, making them hard to detect especially when a detector has little priori knowledge of their traffic features. In this article, we propose a generic and sensitive detection approach, which can simultaneously (i) identify various types of channels without their traffic knowledge and (ii) maintain reasonable performance on small traffic samples. The basis of our approach is the finding that the short-term timing behavior of covert and legitimate traffic is significantly different from the perspective of inter-packet delays’ variation. This phenomenon can be a generic reference to detect various channels because it is resistant to major channel disguising techniques which only mimic long-term traffic features, while it is also a sensitive reference to spot small-volume covert transmission since it can capture traffic anomalies in a fine-grained manner. To obtain the inner patterns of inter-packet delays’ variation, we design a context-sensitive feature-extraction technique. This technique transforms each raw inter-packet delay into a discrete counterpart based on its contextual properties, thus extracting its variation features and reducing traffic data complexity. Then we learn legitimate variation patterns using a neural network model, and identify samples showing anomalous variation as covert. The experimental results show that our approach effectively detects all currently representative channels in the absence of their knowledge, presenting once to twice higher sensitivity than the state-of-the-art solutions. Haozhi Li, Tian Song 0004, Yating Yang |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2022 | Energy-Efficient Cooperative Caching for Information-Centric Wireless Sensor NetworkingabstractAlthough information-centric networking (ICN) can boost the performance of sensor data delivery in wireless sensor networks (WSNs), its energy efficiency may be significantly impaired due to the receiver-driven request-response transmission model. ICN sensors have to keep long-time active mode to wait for the requests that may arrive at any time, resulting nontrivial energy waste for energy-constrained WSNs. An intuitive method to conserve energy is to schedule partial nodes to fall asleep, but at the cost of extra data fetching delay. In this work, we investigate multihop cooperative caching for green WSN over receiver-driven ICNs. By leveraging regional caches, not only can the energy be saved, but also the data fetching delay can be reduced. We formulate those two cooperation gains and capture the tradeoff between energy saving and delay reduction. Based on the optimal tradeoff, we present a solution of scoped multihop cooperative caching and elaborate a green cooperation policy for each sensor to determine its operation scope and conduct a proper cooperation decision. Our simulation results show that our proposed multihop cooperative caching can save more than 91.0% energy and achieve 28.9% delay reduction than the basic ICWSN, and also reduce 60.2% energy consumption than the existing cooperative caching schemes. Yating Yang, Tian Song 0004 |
IEEE Internet Things J. | 1 |
| 2021 | Towards reliable and efficient data retrieving in ICN-based satellite networks
Yating Yang, Tian Song 0004, Weijia Yuan, Jianping An |
J. Netw. Comput. Appl. | 1 |
| 2019 | A Multi-View Spatial-Temporal Network for Vehicle Refueling Demand Inference
Bo Ma 0004, Yating Yang |
KSEM (2) | 2 |
| 2019 | Research of Uyghur-Chinese Machine Translation System Combination Based on Semantic Information
Xiao Li 0007, Yating Yang, Azmat Anwar, Rui Dong 0002 |
NLPCC (2) | 3 |
| 2019 | OpenCache: A lightweight regional cache collaboration approach in hierarchical-named ICN
Yating Yang, Beichuan Zhang 0001 |
Comput. Commun. | 1 |
| 2019 | Convolutional Attention Networks for Scene Text RecognitionabstractIn this article, we present Convoluitional Attention Networks (CAN) for unconstrained scene text recognition. Recent dominant approaches for scene text recognition are mainly based on Convolutional Neural Networks (CNN) and Recurrent Neural Networks (RNN), where the CNN encodes images and the RNN generates character sequences. Our CAN is different from these methods; our CAN is completely built on CNN and includes an attention mechanism. The distinctive characteristics of our method include (i) CAN follows encoder-decoder architecture, in which the encoder is a deep two-dimensional CNN and the decoder is a one-dimensional CNN; (ii) the attention mechanism is applied in every convolutional layer of the decoder, and we propose a novel spatial attention method using average pooling; and (iii) position embeddings are equipped in both a spatial encoder and a sequence decoder to give our networks a sense of location. We conduct experiments on standard datasets for scene text recognition, including Street View Text , IIIT5K, and ICDAR datasets. The experimental results validate the effectiveness of different components and show that our convolutional-based method achieves state-of-the-art or competitive performance over prior works, even without the use of RNN. Hongtao Xie 0001, Shancheng Fang, Zhengjun Zha, Yating Yang, Yan Li 0068, Yongdong Zhang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2018 | Toward Better Loanword Identification in Uyghur Using Cross-lingual Word EmbeddingsabstractTo enrich vocabulary of low resource settings, we proposed a novel method which identify loanwords in monolingual corpora. More specifically, we first use cross-lingual word embeddings as the core feature to generate semantically related candidates based on comparable corpora and a small bilingual lexicon; then, a log-linear model which combines several shallow features such as pronunciation similarity and hybrid language model features to predict the final results. In this paper, we use Uyghur as the receipt language and try to detect loanwords in four donor languages: Arabic, Chinese, Persian and Russian. We conduct two groups of experiments to evaluate the effectiveness of our proposed approach: loanword identification and OOV translation in four language pairs and eight translation directions (Uyghur-Arabic, Arabic-Uyghur, Uyghur-Chinese, Chinese-Uyghur, Uyghur-Persian, Persian-Uyghur, Uyghur-Russian, and Russian-Uyghur). Experimental results on loanword identification show that our method outperforms other baseline models significantly. Neural machine translation models integrating results of loanword identification experiments achieve the best results on OOV translation(with 0.5-0.9 BLEU improvements) Chenggang Mi 0002, Yating Yang, Lei Wang 0065, Xi Zhou 0007, Tonghai Jiang |
COLING | 2 |
| 2018 | Improved Spoken Uyghur Segmentation for Neural Machine TranslationabstractTo increase vocabulary overlap in spoken Uyghur neural machine translation (NMT), we propose a novel method to enhance the common used subword units based segmentation method. In particular, we apply a log-linear model as the main framework and integrate several features such as subword, morphological information, bilingual word alignment and monolingual language model into it. Experimental results show that spoken Uyghur segmentation with our proposed method improves the performance of the spoken Uyghur-Chinese NMT significantly (yield up to 1.52 BLEU improvements). Chenggang Mi 0002, Yating Yang, Xi Zhou 0007, Lei Wang 0065, Tonghai Jiang |
ICTAI | 2 |
| 2018 | A Neural Network Based Model for Loanword Identification in Uyghur
Chenggang Mi 0002, Yating Yang, Lei Wang 0065, Xi Zhou 0007, Tonghai Jiang |
LREC | 2 |
| 2017 | RwHash: Rewritable Hash Table for Fast Network Processing with Dynamic Membership UpdatesabstractHash table is one of the most fundamental and critical data structures for membership query and maintenance. However, the performance of a standard hash table degrades greatly when the hash collision is large due to high load factor or unpredictable dynamic membership updates, especially per-packet updates in network processing. In this paper, we shape a hash table from the conventional slim-and-tall style to a wide-and-short style by facilitating an extension of logical cache block. Then, a cache aware hash table (CaHash) is given and explored in detail. Based on an observation that the operation sequences may be in a potential and probabilistic successive order, especially for network applications, a rewritable hash table (RwHash) is finally proposed, which provides two rewritable policies to dynamically move elements within a bucket when updating. Theoretical analysis shows that, no matter what load factor and collision are, RwHash can achieve near-optimal performance the same as the performance when a standard hash table in the case of no collision. Real experiments show that RwHash can achieve 4.10 times speedup in some parameter and even more with different configurations than a standard hash table in the case of heavy collisions. Our approaches are elegantly practical in implementation for both software and hardware. Yating Yang, Patrick Crowley |
ANCS | 2 |
| 2017 | On the Feasibility of Inter-Domain Routing via a Small Broker SetabstractThe current inter-domain routing protocol, namely, the Border Gateway Protocol (BGP), cannot provide end-to-end (E2E) quality-of-service (QoS) guarantees. The main reason is that an autonomous system (AS) can only receive guarantees from its first hop ASes via service level agreements (SLAs). But beyond the first hop, QoS along the path from source to destination AS is not within the source AS's control regime. In this paper, we investigate the feasibility of providing high QoS-guaranteed E2E transit services by utilizing a (small) set of ASes/IXPs to serve as "brokers" to provide supervision, control and resource negotiation. Finding an optimal set of ASes as brokers can be formulated as a Maximum Coverage with B-dominating path Guarantee (MCBG) problem, which we prove to be NP-hard. To address this problem, we design a (1-e-1/4)-approximation algorithm and also an efficient heuristic algorithm when considering additional constraints (e.g., path length). Based on the current Internet topology, we discover a "3540-alliance" subset (accounting only 6.8%) of 52,079 ASes/IXPs, which can provide high QoS guarantees for 99.29% E2E connections. Dong Lin, David Shui Wing Hui, Weijie Wu, Tingwei Liu, Yating Yang, Yi Wang 0004, John C. S. Lui, Gong Zhang 0001 |
ICDCS | 5 |
| 2017 | Learning Bilingual Lexicon for Low-Resource Language Pairs
ShaoLin Zhu, Xiao Li 0007, Yating Yang, Lei Wang 0065, Chenggang Mi 0002 |
NLPCC | 3 |
| 2016 | A Bilingual Discourse Corpus and Its Applications
Yang Liu 0085, Jiajun Zhang 0001, Chengqing Zong, Yating Yang, Xi Zhou 0007 |
LREC | 4 |
| 2016 | Recurrent Neural Network Based Loanwords Identification in Uyghur
Chenggang Mi 0002, Yating Yang, Xi Zhou 0007, Lei Wang 0065, Xiao Li 0007, Tonghai Jiang |
PACLIC | 2 |
| 2016 | A Semisupervised Tag-Transition-Based Markovian Model for Uyghur Morphology AnalysisabstractMorphological analysis, which includes analysis of part-of-speech (POS) tagging, stemming, and morpheme segmentation, is one of the key components in natural language processing (NLP), particularly for agglutinative languages. In this article, we investigate the morphological analysis of the Uyghur language, which is the native language of the people in the Xinjiang Uyghur autonomous region of western China. Morphological analysis of Uyghur is challenging primarily because of factors such as (1) ambiguities arising due to the likelihood of association of a multiple number of POS tags with a word stem or a multiple number of functional tags with a word suffix, (2) ambiguous morpheme boundaries, and (3) complex morphopholonogy of the language. Further, the unavailability of a manually annotated training set in the Uyghur language for the purpose of word segmentation makes Uyghur morphological analysis more difficult. In our proposed work, we address these challenges by undertaking a semisupervised approach of learning a Markov model with the help of a manually constructed dictionary of “suffix to tag” mappings in order to predict the most likely tag transitions in the Uyghur morpheme sequence. Due to the linguistic characteristics of Uyghur, we incorporate a prior belief in our model for favoring word segmentations with a lower number of morpheme units. Empirical evaluation of our proposed model shows an accuracy of about 82%. We further improve the effectiveness of the tag transition model with an active learning paradigm. In particular, we manually investigated a subset of words for which the model prediction ambiguity was within the top 20%. Manually incorporating rules to handle these erroneous cases resulted in an overall accuracy of 93.81%. Eziz Tursun, Debasis Ganguly, Osman Turghun, Yating Yang, Ghalip Abdukerim, Junlin Zhou, Qun Liu 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 4 |
| 2015 | Optimized Uyghur Segmentation for Statistical Machine Translation
Chenggang Mi 0002, Yating Yang, Rui Dong 0002, Xi Zhou 0007, Lei Wang 0065, Xiao Li 0007, Tonghai Jiang, Osman Turghun |
NLDB | 2 |
| 2014 | Detection of Loan Words in Uyghur Texts
Chenggang Mi 0002, Yating Yang, Lei Wang 0065, Xiao Li 0007, Kamali Dalielihan |
NLPCC | 2 |
| 2013 | Discriminative Learning with Natural Annotations: Word Segmentation as a Case Study
Wenbin Jiang 0002, Yajuan Lü, Yating Yang, Qun Liu 0001 |
ACL (1) | 4 |