Xiaobin Wang

dblp:17/5812 · DBLP profile ↗
← Back
42ranked-venue papers
5as first author
24since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 24 · 3 first-author · 18 since 2021Systems, architecture and hardware · 13 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Addressing the Computational Divide in Vehicular Networks: A Lifecycle-Aware Task Offloading Framework With Deep Reinforcement Learning
abstract
The Internet of Vehicles (IoV), a critical large-scale Internet of Things (IoT) application, faces a fundamental sustainability challenge stemming from the lifecycle mismatch between long-duration vehicular hardware and rapidly evolving software. This mismatch creates a widening “computational divide,” where aging vehicles with limited onboard resources cannot support modern data-intensive applications, thereby fragmenting the ecosystem and undermining its collective intelligence. To address this systemic issue, this paper proposes a novel lifecycle-aware task offloading framework. Instead of treating vehicular heterogeneity as a static liability, our framework transforms it into a dynamic, cooperative resource-sharing opportunity. At the core of this framework is a Deep Reinforcement Learning (DRL) agent deployed on resource-constrained vehicles, which orchestrates task offloading decisions to co-optimize for latency and energy consumption. A key innovation is the “Performance Capacity” metric, a multi-dimensional and predictive state assessment mechanism. This metric enables informed decision-making by intelligently fusing a vehicle’s static hardware profile, its dynamically predicted future workload, and its cooperation reputation. Furthermore, an integrated credit-based incentive mechanism is designed to ensure the long-term economic viability of this cooperative ecosystem. Simulation results demonstrate that our framework significantly reduces task latency and energy consumption, particularly under high-load conditions, offering a scalable and sustainable solution for the continuous evolution of heterogeneous IoV systems.
Xiaobin Wang, Yuntao Zou, Zeling Xu, Wei Wang 0077, Dagang Li 0001
IEEE Internet Things J.1
2026 Goal programming-based Wasserstein distributionally robust equilibrium vehicle routing design with carbon cost intensity
Fanghao Yin, Xiaobin Wang
Inf. Sci.3
2025 Agentic Knowledgeable Self-awareness
abstract
Shuofei Qiao, Zhisong Qiu, Baochang Ren, Xiaobin Wang, Xiangyuan Ru, Ningyu Zhang, Xiang Chen, Yong Jiang, Pengjun Xie, Fei Huang, Huajun Chen. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Shuofei Qiao, Zhisong Qiu, Baochang Ren, Xiaobin Wang, Xiangyuan Ru, Ningyu Zhang 0001, Xiang Chen 0016, Yong Jiang 0005, Pengjun Xie, Fei Huang 0002, Huajun Chen
ACL (1)4
2025 Benchmarking Agentic Workflow Generation
abstract
Large Language Models (LLMs), with their exceptional ability to handle a wide range of tasks, have driven significant advancements in tackling reasoning and planning tasks, wherein decomposing complex problems into executable workflows is a crucial step in this process. Existing workflow evaluation frameworks either focus solely on holistic performance or suffer from limitations such as restricted scenario coverage, simplistic workflow structures, and lax evaluation standards. To this end, we introduce WorfBench, a unified workflow generation benchmark with multi-faceted scenarios and intricate graph workflow structures. Additionally, we present WorfEval, a systemic evaluation protocol utilizing subsequence and subgraph matching algorithms to accurately quantify the LLM agent's workflow generation capabilities. Through comprehensive evaluations across different types of LLMs, we discover distinct gaps between the sequence planning capabilities and graph planning capabilities of LLM agents, with even GPT-4 exhibiting a gap of around 15%. We also train two open-source models and evaluate their generalization abilities on held-out tasks. Furthermore, we observe that the generated workflows can enhance downstream tasks, enabling them to achieve superior performance with less time during inference. Code and dataset are available at https://github.com/zjunlp/WorfBench.
Shuofei Qiao, Runnan Fang, Zhisong Qiu, Xiaobin Wang, Ningyu Zhang 0001, Yong Jiang 0005, Pengjun Xie, Fei Huang 0002, Huajun Chen
ICLR4
2024 Three Heads Are Better than One: Improving Cross-Domain NER with Progressive Decomposed Network
abstract
Cross-domain named entity recognition (NER) tasks encourage NER models to transfer knowledge from data-rich source domains to sparsely labeled target domains. Previous works adopt the paradigms of pre-training on the source domain followed by fine-tuning on the target domain. However, these works ignore that general labeled NER source domain data can be easily retrieved in the real world, and soliciting more source domains could bring more benefits. Unfortunately, previous paradigms cannot efficiently transfer knowledge from multiple source domains. In this work, to transfer multiple source domains' knowledge, we decouple the NER task into the pipeline tasks of mention detection and entity typing, where the mention detection unifies the training object across domains, thus providing the entity typing with higher-quality entity mentions. Additionally, we request multiple general source domain models to suggest the potential named entities for sentences in the target domain explicitly, and transfer their knowledge to the target domain models through the knowledge progressive networks implicitly. Furthermore, we propose two methods to analyze in which source domain knowledge transfer occurs, thus helping us judge which source domain brings the greatest benefit. In our experiment, we develop a Chinese cross-domain NER dataset. Our model improved the F1 score by an average of 12.50% across 8 Chinese and English datasets compared to models without source domain data.
Xuming Hu, Zhaochen Hong, Yong Jiang 0005, Zhichao Lin, Xiaobin Wang, Pengjun Xie, Philip S. Yu
AAAI5
2024 EcomGPT: Instruction-Tuning Large Language Models with Chain-of-Task Tasks for E-commerce
abstract
Recently, instruction-following Large Language Models (LLMs) , represented by ChatGPT, have exhibited exceptional performance in general Natural Language Processing (NLP) tasks. However, the unique characteristics of E-commerce data pose significant challenges to general LLMs. An LLM tailored specifically for E-commerce scenarios, possessing robust cross-dataset/task generalization capabilities, is a pressing necessity. To solve this issue, in this work, we proposed the first E-commerce instruction dataset EcomInstruct, with a total of 2.5 million instruction data. EcomInstruct scales up the data size and task diversity by constructing atomic tasks with E-commerce basic data types, such as product information, user reviews. Atomic tasks are defined as intermediate tasks implicitly involved in solving a final task, which we also call Chain-of-Task tasks. We developed EcomGPT with different parameter scales by training the backbone model BLOOMZ with the EcomInstruct. Benefiting from the fundamental semantic understanding capabilities acquired from the Chain-of-Task tasks, EcomGPT exhibits excellent zero-shot generalization capabilities. Extensive experiments and human evaluations demonstrate that EcomGPT outperforms ChatGPT in term of cross-dataset/task generalization on E-commerce tasks. The EcomGPT will be public at https://github.com/Alibaba-NLP/EcomGPT.
Yangning Li, Shirong Ma, Xiaobin Wang, Shen Huang, Chengyue Jiang, Hai-Tao Zheng 0002, Pengjun Xie, Fei Huang 0002, Yong Jiang 0005
AAAI3
2024 SeqGPT: An Out-of-the-Box Large Language Model for Open Domain Sequence Understanding
abstract
Large language models (LLMs) have shown impressive abilities for open-domain NLP tasks. However, LLMs are sometimes too footloose for natural language understanding (NLU) tasks which always have restricted output and input format. Their performances on NLU tasks are highly related to prompts or demonstrations and are shown to be poor at performing several representative NLU tasks, such as event extraction and entity typing. To this end, we present SeqGPT, a bilingual (i.e., English and Chinese) open-source autoregressive model specially enhanced for open-domain natural language understanding. We express all NLU tasks with two atomic tasks, which define fixed instructions to restrict the input and output format but still ``open'' for arbitrarily varied label sets. The model is first instruction-tuned with extremely fine-grained labeled data synthesized by ChatGPT and then further fine-tuned by 233 different atomic tasks from 152 datasets across various domains. The experimental results show that SeqGPT has decent classification and extraction ability, and is capable of performing language understanding tasks on unseen domains. We also conduct empirical studies on the scaling of data and model size as well as on the transfer across tasks. Our models are accessible at https://github.com/Alibaba-NLP/SeqGPT.
Tianyu Yu 0002, Chengyue Jiang, Chao Lou, Shen Huang, Xiaobin Wang, Wei Liu 0131, Jiong Cai, Yangning Li, Kewei Tu, Hai-Tao Zheng 0002, Ningyu Zhang 0001, Pengjun Xie, Fei Huang 0002, Yong Jiang 0005
AAAI5
2024 Effective Demonstration Annotation for In-Context Learning via Language Model-Based Determinantal Point Process
abstract
In-context learning (ICL) is a few-shot learning paradigm that involves learning mappings through input-output pairs and appropriately applying them to new instances.Despite the remarkable ICL capabilities demonstrated by Large Language Models (LLMs), existing works are highly dependent on large-scale labeled support sets, not always feasible in practical scenarios.To refine this approach, we focus primarily on an innovative selective annotation mechanism, which precedes the standard demonstration retrieval.We introduce the Language Model-based Determinant Point Process (LM-DPP) that simultaneously considers the uncertainty and diversity of unlabeled instances for optimal selection.Consequently, this yields a subset for annotation that strikes a trade-off between the two factors.We apply LM-DPP to various language models, including GPT-J, LlaMA, and GPT-3.Experimental results on 9 NLU and 2 Generation datasets demonstrate that LM-DPP can effectively select canonical examples.Further analysis reveals that LLMs benefit most significantly from subsets that are both low uncertainty and high diversity.
Peng Wang 0104, Xiaobin Wang, Chao Lou, Shengyu Mao, Pengjun Xie, Yong Jiang 0005
EMNLP2
2024 Compact modeling of layout parasitic effects on power MOSFET switching
abstract
This paper studies the effect of layout parasitics on the switching behavior of power MOSFETs using a distributed modeling approach. It uses Verilog-A compact models for individual transistor cells to access the impact of layout parasitics on device performance. Parameter values of the SPICE model are validated through comparisons with a TCAD model of a silicon split-gate trench MOSFET. The compact model for the distributed MOSFET network enables accurate analysis of gate signal delay times, in contrast to the conventional approach employing constant RC networks. The modeling technique is also used to investigate the internal current and power distribution across the device layout during turn-on and turn-off switching events.
Xiaobin Wang
IECON3
2024 High-dimensional quantile mediation analysis with application to a birth cohort study of mother-newborn pairs
abstract
MOTIVATION: There has been substantial recent interest in developing methodology for high-dimensional mediation analysis. Yet, the majority of mediation statistical methods lean heavily on mean regression, which limits their ability to fully capture the complex mediating effects across the outcome distribution. To bridge this gap, we propose a novel approach for selecting and testing mediators throughout the full range of the outcome distribution spectrum. RESULTS: The proposed high-dimensional quantile mediation model provides a comprehensive insight into how potential mediators impact outcomes via their mediation pathways. This method's efficacy is demonstrated through extensive simulations. The study presents a real-world data application examining the mediating effects of DNA methylation on the relationship between maternal smoking and offspring birthweight. AVAILABILITY AND IMPLEMENTATION: Our method offers a publicly available and user-friendly function qHIMA(), which can be accessed through the R package HIMA at https://CRAN.R-project.org/package=HIMA.
Xiumei Hong, Yinan Zheng, Lifang Hou, Cheng Zheng 0005, Xiaobin Wang, Lei Liu 0004
Bioinform.6
2023 Exploring Lottery Prompts for Pre-trained Language Models
abstract
Yulin Chen, Ning Ding, Xiaobin Wang, Shengding Hu, Haitao Zheng, Zhiyuan Liu, Pengjun Xie. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Yulin Chen 0001, Ning Ding 0002, Xiaobin Wang, Shengding Hu, Hai-Tao Zheng 0002, Zhiyuan Liu 0001, Pengjun Xie
ACL (1)3
2023 MANNER: A Variational Memory-Augmented Model for Cross Domain Few-Shot Named Entity Recognition
abstract
This paper focuses on the task of cross domain few-shot named entity recognition (NER), which aims to adapt the knowledge learned from source domain to recognize named entities in target domain with only a few labeled examples.To address this challenging task, we propose MANNER, a variational memoryaugmented few-shot NER model.Specifically, MANNER uses a memory module to store information from the source domain and then retrieve relevant information from the memory to augment few-shot tasks in the target domain.In order to effectively utilize the information from memory, MANNER uses optimal transport to retrieve and process information from memory, which can explicitly adapt the retrieved information from source domain to target domain and improve the performance in the cross domain few-shot setting.We conduct experiments on both English and Chinese cross domain fewshot NER datasets, and the experimental results demonstrate that MANNER can achieve superior performance 1 .
Jinyuan Fang, Xiaobin Wang, Zaiqiao Meng, Pengjun Xie, Fei Huang 0002, Yong Jiang 0005
ACL (1)2
2023 Recall, Expand, and Multi-Candidate Cross-Encode: Fast and Accurate Ultra-Fine Entity Typing
abstract
Ultra-fine entity typing (UFET) predicts extremely free-formed types (e.g., president, politician) of a given entity mention (e.g., Joe Biden) in context.State-of-the-art (SOTA) methods use the cross-encoder (CE) based architecture.CE concatenates a mention (and its context) with each type and feeds the pair into a pretrained language model (PLM) to score their relevance.It brings deeper interaction between the mention and the type to reach better performance but has to perform N (the type set size) forward passes to infer all the types of a single mention.CE is therefore very slow in inference when the type set is large (e.g., N = 10k for UFET).To this end, we propose to perform entity typing in a recall-expand-filter manner.The recall and expansion stages prune the large type set and generate K (typically much smaller than N ) most relevant type candidates for each mention.At the filter stage, we use a novel model called MCCE to concurrently encode and score all these K candidates in only one forward pass to obtain the final type prediction.We investigate different model options for each stage and conduct extensive experiments to compare each option, experiments show that our method reaches SOTA performance on UFET and is thousands of times faster than the CE-based architecture.We also found our method is very effective in fine-grained (130 types) and coarse-grained (9 types) entity typing.
Chengyue Jiang, Wenyang Hui, Yong Jiang 0005, Xiaobin Wang, Pengjun Xie, Kewei Tu
ACL (1)4
2023 Dense-and-Similar Object detection in aerial images
Xiaobin Wang, Haohui Sun, Dekang Zhu
Pattern Recognit. Lett.1
2022 Parallel Instance Query Network for Named Entity Recognition
abstract
Yongliang Shen, Xiaobin Wang, Zeqi Tan, Guangwei Xu, Pengjun Xie, Fei Huang, Weiming Lu, Yueting Zhuang. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Yongliang Shen 0001, Xiaobin Wang, Zeqi Tan, Pengjun Xie, Fei Huang 0002, Weiming Lu 0001, Yueting Zhuang
ACL (1)2
2022 Identifying Chinese Opinion Expressions with Extremely-Noisy Crowdsourcing Annotations
abstract
Recent works of opinion expression identification (OEI) rely heavily on the quality and scale of the manually-constructed training corpus, which could be extremely difficult to satisfy.Crowdsourcing is one practical solution for this problem, aiming to create a large-scale but quality-unguaranteed corpus.In this work, we investigate Chinese OEI with extremelynoisy crowdsourcing annotations, constructing a dataset at a very low cost.Following Zhang et al. (2021), we train the annotator-adapter model by regarding all annotations as goldstandard in terms of crowd annotators, and test the model by using a synthetic expert, which is a mixture of all annotators.As this annotatormixture for testing is never modeled explicitly in the training phase, we propose to generate synthetic training samples by a pertinent mixup strategy to make the training and testing highly consistent.The simulation experiments on our constructed dataset show that crowdsourcing is highly promising for OEI, and our proposed annotator-mixup can further enhance the crowdsourcing modeling.
Xin Zhang 0097, Yueheng Sun, Meishan Zhang, Xiaobin Wang, Min Zhang 0005
ACL (1)5
2022 Domain-Specific NER via Retrieving Correlated Samples
abstract
Successful Machine Learning based Named Entity Recognition models could fail on texts from some special domains, for instance, Chinese addresses and e-commerce titles, where requires adequate background knowledge. Such texts are also difficult for human annotators. In fact, we can obtain some potentially helpful information from correlated texts, which have some common entities, to help the text understanding. Then, one can easily reason out the correct answer by referencing correlated samples. In this paper, we suggest enhancing NER models with correlated samples. We draw correlated samples by the sparse BM25 retriever from large-scale in-domain unlabeled data. To explicitly simulate the human reasoning process, we perform a training-free entity type calibrating by majority voting. To capture correlation features in the training stage, we suggest to model correlated samples by the transformer-based multi-instance cross-encoder. Empirical results on datasets of the above two domains show the efficacy of our methods.
Xin Zhang 0097, Yong Jiang 0005, Xiaobin Wang, Xuming Hu, Yueheng Sun, Pengjun Xie, Meishan Zhang
COLING3
2022 AISHELL-NER: Named Entity Recognition from Chinese Speech
abstract
Named Entity Recognition (NER) from speech is among Spoken Language Understanding (SLU) tasks, aiming to extract semantic information from the speech signal. NER from speech is usually made through a two-step pipeline that consists of (1) processing the audio using an Automatic Speech Recognition (ASR) system and (2) applying an NER tagger to the ASR outputs. Recent works have shown the capability of the End-to-End (E2E) approach for NER from English and French speech, which is essentially entity-aware ASR. However, due to the many homophones and polyphones that exist in Chinese, NER from Chinese speech is effectively a more challenging task. In this paper, we introduce a new dataset AISEHLL-NER for NER from Chinese speech. Extensive experiments are conducted to explore the performance of several state-of-the-art methods. The results demonstrate that the performance could be improved by combining entity-aware ASR and pretrained NER tagger, which can be easily applied to the modern SLU pipeline. The dataset is publicly available at github.com/Alibaba-NLP/AISHELL-NER.
Boli Chen, Xiaobin Wang, Pengjun Xie, Meishan Zhang, Fei Huang 0002
ICASSP3
2022 Solving Poker Games Efficiently: Adaptive Memory based Deep Counterfactual Regret Minimization
abstract
Poker game has become one of the most prevailing benchmark environment to discover algorithms for sequential games with imperfect information (SGII). However, in games with large state space, it is hard to traverse the whole game tree. This is because the space of history is exponentially increasing with the input size of the game. Other attempts like truncating the game tree with certain length have also been made to solve this problem. But determine the most suitable length could require enormous amount of resources. All of these obstacles make algorithms for SGII much harder to design. To solve this kind of problem, we propose the adaptive memory sampling method which aims to find the distribution of the sampling length by using posterior sampling to update it iteratively. In the real-world human interaction, to what extent a human memory can last often varies significantly depending on the importance of the interaction trajectory. So we also adopted the Long Short-Term Memory (LSTM) network as the sub-procedure to classify the histories and making prediction of future game states and actions based on historical sampled data. According to our theoretical analysis, our method performs better than the state-of-the-art algorithms. On the other hand, The empirical results support our results.
Shuqing Shi, Xiaobin Wang, Dong Hao, Zhiyou Yang, Hong Qu 0002
IJCNN2
2022 DAMO-NLP at NLPCC-2022 Task 2: Knowledge Enhanced Robust NER for Speech Entity Linking
Shen Huang, Yuchen Zhai, Xinwei Long, Yong Jiang 0005, Xiaobin Wang, Yin Zhang 0006, Pengjun Xie
NLPCC (2)5
2022 Base station computing force resource load balancing strategy for distributed machine learning
abstract
With the emergence of terminal services such as VR, Internet of Vehicles, and autonomous driving that require enormous computing resources and network transmission resources, the computing power and network load of existing 5G base stations have been difficult to bear. Mobile edge computing technology effectively integrates the two technologies of mobile network and Internet, adding computing, storage, data processing and other functions on the mobile network side, which building an open platform to implant applications. The existing technology has shortcomings such as insufficient computing power of communication base stations, limited resources of a single edge computing node, etc., It is difficult to meet the development needs of the industry. Therefore, this paper proposes a base station computing power load balancing method based on distributed machine learning. Through distributed machine learning, the communication base station is used as a computing node, and the idle computing power of the base station is called to achieve a reasonable allocation of the computing power resources of the base station.
Mingkang Song, Mengke Yao, Xiaobin Wang, Tenghui Ke
ICSS3
2021 Few-NERD: A Few-shot Named Entity Recognition Dataset
abstract
Ning Ding, Guangwei Xu, Yulin Chen, Xiaobin Wang, Xu Han, Pengjun Xie, Haitao Zheng, Zhiyuan Liu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Ning Ding 0002, Yulin Chen 0001, Xiaobin Wang, Xu Han 0007, Pengjun Xie, Hai-Tao Zheng 0002, Zhiyuan Liu 0001
ACL/IJCNLP (1)4
2021 Prototypical Representation Learning for Relation Extraction
Ning Ding 0002, Xiaobin Wang, Rui Wang 0005, Pengjun Xie, Ying Shen 0001, Fei Huang 0002, Hai-Tao Zheng 0002, Rui Zhang 0003
ICLR2
2021 Bio-inspired Model Based on Global-Local Hybrid Learning in Spiking Neural Network
abstract
Bringing machines up to human-level visual processing capabilities is an attractive research topic for decades. Deep neural networks (DNNs), inspired by the hierarchical structure of the human primary visual cortex at a macroscopic level, have achieved state-of-the-art performance in many applications. However, their practical applications remain limited due to the requisition of massive computing resources. Spiking neural networks (SNNs) simulate the spike-based information process of the biological neural system from the microscopic view and hold greater potential to ultra-low-power computations. In this paper, we imitate the human visual system from both the micro and macro scales and make the following contributions: (1) Inspired by the lateral effect between real neurons, we propose a Global-Local Hybrid Spike-Timing-Dependent Plasticity (GLHSTDP) algorithm that combines STDP with lateral synaptic learning mechanism, to train the spiking neural network. (2) We construct a deep spiking neural network (DSNN) to mimic the visual information processing mechanism in the human brain. Experimental results demonstrate that the proposed DSNN model equipped with the proposed learning algorithm works in a totally spike-based manner and achieve competitive accuracies on both the Caltech 101 and the MNIST datasets.
Xiaobin Wang, Hong Qu 0002, Yi Chen 0034, Xiaoling Luo 0001
IJCNN2
2020 Coupling Distant Annotation and Adversarial Training for Cross-Domain Chinese Word Segmentation
abstract
Fully supervised neural approaches have achieved significant progress in the task of Chinese word segmentation (CWS).Nevertheless, the performance of supervised models tends to drop dramatically when they are applied to outof-domain data.Performance degradation is caused by the distribution gap across domains and the out of vocabulary (OOV) problem.In order to simultaneously alleviate these two issues, this paper proposes to couple distant annotation and adversarial training for crossdomain CWS.For distant annotation, we rethink the essence of "Chinese words" and design an automatic distant annotation mechanism that does not need any supervision or pre-defined dictionaries from the target domain.The approach could effectively explore domain-specific words and distantly annotate the raw texts for the target domain.For adversarial training, we develop a sentence-level training procedure to perform noise reduction and maximum utilization of the source domain information.Experiments on multiple realworld datasets across various domains show the superiority and robustness of our model, significantly outperforming previous state-ofthe-art cross-domain CWS methods.
Ning Ding 0002, Dingkun Long, Muhua Zhu, Pengjun Xie, Xiaobin Wang, Hai-Tao Zheng 0002
ACL6
2020 Green supply chain analysis under cost sharing contract with uncertain information based on confidence level
Nana Ma, Rong Gao 0003, Xiaobin Wang
Soft Comput.3
2019 Unsupervised Learning Helps Supervised Neural Word Segmentation
abstract
By exploiting unlabeled data for further performance improvement for Chinese word segmentation, this work makes the first attempt at exploring adding unsupervised segmentation information into neural supervised segmenter. We survey various effective strategies, including extending the character embedding, augmenting the word score and applying multi-task learning, for leveraging unsupervised information derived from abundant unlabeled data. Experiments on standard data sets show that the explored strategies indeed improve the recall rate of out-of-vocabulary words and thus boost the segmentation accuracy. Moreover, the model enhanced by the proposed methods outperforms state-of-theart models in closed test and shows promising improvement trend when adopting three different strategies with the help of a large unlabeled data set. Our thorough empirical study eventually verifies the proposed approach outperforms the widelyused pre-training approach in terms of effectively making use of freely abundant unlabeled data.
Xiaobin Wang, Deng Cai 0002, Linlin Li 0001, Hai Zhao 0001, Luo Si
AAAI1
2019 Learning Models from Data with Measurement Error: Tackling Underreporting
abstract
Measurement error in observational datasets can lead to systematic bias in inferences based on these datasets. As studies based on observational data are increasingly used to inform decisions with real-world impact, it is critical that we develop a robust set of techniques for analyzing and adjusting for these biases. In this paper we present a method for estimating the distribution of an outcome given a binary exposure that is subject to underreporting. Our method is based on a missing data view of the measurement error problem, where the true exposure is treated as a latent variable that is marginalized out of a joint model. We prove three different conditions under which the outcome distribution can still be identified from data containing only error-prone observations of the exposure. We demonstrate this method on synthetic data and analyze its sensitivity to near violations of the identifiability conditions. Finally, we use this method to estimate the effects of maternal smoking and heroin use during pregnancy on childhood obesity, two import problems from public health. Using the proposed method, we estimate these effects using only subject-reported drug use data and refine the range of estimates generated by a sensitivity analysis-based approach. Further, the estimates produced by our method are consistent with existing literature on both the effects of maternal smoking and the rate at which subjects underreport smoking.
Roy Adams, Yuelong Ji, Xiaobin Wang, Suchi Saria
ICML3
2017 Persistent and Nonpersistent Error Optimization for STT-RAM Cell Design
abstract
Rapidly increasing demands for memory capacity and severe technical scaling challenges of conventional memory technologies motivated recent investments on next-generation nonvolatile memory technologies. As a promising candidate, spin-transfer torque random access memory (STT-RAM) has demonstrated many attractive properties, such as nanosecond access time, high integration density, nonvolatility, and excellent CMOS integration compatibility. However, similar to all other nano-devices, the performance and reliability of STT-RAM cells are greatly affected by process variations, device operating uncertainties, and environmental fluctuations. As a result, the read and write operations of STT-RAM demonstrate some variabilities and errors. In this paper, we systematically analyze the impacts of CMOS and magnetic tunneling junction (MTJ) process variations, MTJ resistance switching randomness that are induced by intrinsic thermal fluctuations, and working temperature changes on STT-RAM cell designs. The STT-RAM cell reliability issues in both read and write operations are first investigated. A combined circuit and magnetic simulation platform is then established to quantitatively study the persistent and nonpersistent errors in STT-RAM cell operations. Our analysis proved the importance of a full statistical design method in STT-RAM designs for design pessimism minimization.
Yaojun Zhang, Bonan Yan, Xiaobin Wang, Yiran Chen 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2015 Biography-Dependent Collaborative Entity Archiving for Slot Filling
abstract
Knowledge Base Population (KBP) tasks, such as slot filling, show the particular importance of entity-oriented automatic relevant document acquisition.Rich, diverse and reliable relevant documents satisfy the fundamental requirement that a KBP system explores the nature of an entity.Towards the bottleneck problem between comprehensiveness and definiteness of acquisition, we propose a collaborative archiving method.In particular we introduce topic modeling methodologies into entity biography profiling, so as to build a bridge between fuzzy and exact matching.On one side, we employ the topics in a small-scale high-quality relevant documents (i.e., exact matching results) to summarize the life slices of a target entity (i.e., biography), and on the other side, we use the biography as a reliable reference material to detect new truly relevant documents from a large-scale partially complete pseudo-feedback (i.e., fuzzy matching results).We leverage the archiving method to enhance slot filling systems.Experiments on KBP corpus show significant improvement over stateof-the-art.
Xiaobin Wang, Tongtao Zhang, Heng Ji 0001
EMNLP2
2012 Asymmetry of MTJ switching and its implication to STT-RAM designs
abstract
As one promising candidate for next-generation nonvolatile memory technologies, spin-transfer torque random access memory (STT-RAM) has demonstrated many attractive features, such as nanosecond access time, high integration density, non-volatility, and good CMOS process compatibility. In this paper, we reveal an important fact that has been neglected in STT-RAM designs for long: the write operation of a STT-RAM cell is asymmetric based on the switching direction of the MTJ (magnetic tunneling junction) device: the mean and the deviation of the write latency for the switching from low- to high-resistance state is much longer or larger than that of the opposite switching. Some special design concerns, e.g., the write-pattern-dependent write reliability, are raised by this observation. We systematically analyze the root reasons to form the asymmetric switching of the MTJ and study their impacts on STT-RAM write operations. These factors include the thermal-induced statistical MTJ magnetization process, asymmetric biasing conditions of NMOS transistors, and both NMOS and MTJ device variations. We also explore the design space of different design methodologies on capturing the switching asymmetry of different STT-RAM cell structures. Our experiment results proved the importance of full statistical design method in STT-RAM designs for design pessimism minimization.
Yaojun Zhang, Xiaobin Wang, Yong Li 0009, Alex K. Jones, Yiran Chen 0001
DATE2
2012 Voltage Driven Nondestructive Self-Reference Sensing Scheme of Spin-Transfer Torque Memory
abstract
Spin-transfer torque random access memory (STT-RAM) has demonstrated great potentials as a universal memory for its fast access speed, zero standby power, excellent scalability, and simplicity of cell structure. However, large process variations of both magnetic tunneling junction (MTJ) and CMOS process severely limit the yield of STT-RAM chips and prevent the massive production from happening. In this paper, we analyze and compare the impacts of process variations on various sensing schemes of STT-RAM design. On top of it, we propose a novel voltage-driven nondestructive self-reference sensing scheme to enhance the STT-RAM chip yield by significantly improving sense margin. Monte Carlo simulations of a 16-Kb STT-RAM array shows that our proposed scheme can achieve the same yield as the previous nondestructive self-reference sensing scheme while improving the sense margin by five times with the similar access performance and power.
Zhenyu Sun 0001, Hai Li 0001, Yiran Chen 0001, Xiaobin Wang
IEEE Trans. Very Large Scale Integr. Syst.4
2011 Geometry variations analysis of TiO2 thin-film and spintronic memristors
abstract
The fourth passive circuit element, memristor, has attracted increased attentions since the first real device was discovered by HP Lab in 2008. Its distinctive characteristic to record the historic profile of the voltage/current through itself creates great potentials in future system design. However, as a nano-scale device, memristor is facing great challenge on process variation control in the manufacturing. In this work, we analyze the impact of the geometry variations on the electrical properties of both TiO2thin-film and spintronic memristors, including line edge roughness and thickness fluctuation. A simple algorithm was proposed to generate a large volume of geometry variation-aware three-dimensional device structures for Monte-Carlo simulations. Our simulation results show that due to the different physical mechanisms, TiO2thin-film memristor and spintronic memristor demonstrate very different electrical characteristics even when exposing them to the same excitations and under the same process variation conditions.
Miao Hu 0002, Hai Li 0001, Yiran Chen 0001, Xiaobin Wang, Robinson E. Pino
ASP-DAC4
2011 STT-RAM cell design optimization for persistent and non-persistent error rate reduction: A statistical design view
abstract
The rapidly increased demands for memory in electronic industry and the significant technical scaling challenges of all conventional memory technologies motivated the researches on the next generation memory technology. As one promising candidate, spin-transfer torque random access memory (STT-RAM) features fast access time, high density, non-volatility, and good CMOS process compatibility. However, like all other nano-scale devices, the performance and reliability of STT-RAM cells are severely affected by process variations, intrinsic device operating uncertainties and environmental fluctuations. In this work, we systematically analyze the impacts of CMOS and MTJ process variations, MTJ switching uncertainties induced by thermal fluctuations and working temperature on the performance and reliability of STT-RAM cells. A combined circuit and magnetic simulation platform is also established to quantitatively analyze the persistent and non-persistent error rates during the STT-RAM cell operations. Finally, an optimization flow and its effectiveness are depicted by using some STT-RAM cell designs as case study.
Yaojun Zhang, Xiaobin Wang, Yiran Chen 0001
ICCAD2
2011 Continuous review inventory model with variable lead time in a fuzzy random environment
Xiaobin Wang
Expert Syst. Appl.1
2011 Design of Last-Level On-Chip Cache Using Spin-Torque Transfer RAM (STT RAM)
abstract
Because of its high storage density with superior scalability, low integration cost and reasonably high access speed, spin-torque transfer random access memory (STT RAM) appears to have a promising potential to replace SRAM as last-level on-chip cache (e.g., L2 or L3 cache) for microprocessors. Due to unique operational characteristics of its storage device magnetic tunneling junction (MTJ), STT RAM is inherently subject to a write latency versus read latency tradeoff that is determined by the memory cell size. This paper first quantitatively studies how different memory cell sizing may impact the overall computing system performance, and shows that different computing workloads may have conflicting expectations on memory cell sizing. Leveraging MTJ device switching characteristics, we further propose an STT RAM architecture design method that can make STT RAM cache with relatively small memory cell size perform well over a wide spectrum of computing benchmarks. This has been well demonstrated using CACTI-based memory modeling and computing system performance simulations using SimpleScalar. Moreover, we show that this design method can also reduce STT RAM cache energy consumption by up to 30% over a variety of benchmarks.
Wei Xu 0021, Hongbin Sun 0001, Xiaobin Wang, Yiran Chen 0001, Tong Zhang 0002
IEEE Trans. Very Large Scale Integr. Syst.3
2010 A nondestructive self-reference scheme for Spin-Transfer Torque Random Access Memory (STT-RAM)
abstract
We proposed a novel self-reference sensing scheme for Spin-Transfer Torque Random Access Memory (STT-RAM) to overcome the large bit-to-bit variation of Magnetic Tunneling Junction (MTJ) resistance. Different from all the existing schemes, our solution is nondestructive: The stored value in the STT-RAM cell does NOT need to be overwritten by a reference value. And hence, long write-back operation (of the original stored value) is eliminated. The robustness analyses of the existing scheme and our proposed nondestructive scheme are also presented. The measurement results from a 16kb testing chip successfully confirmed the effectiveness of our technique.
Yiran Chen 0001, Hai Li 0001, Xiaobin Wang, Wenzhong Zhu, Wei Xu 0021, Tong Zhang 0002
DATE3
2010 Spintronic memristor devices and application
abstract
Spintronic memristor devices based upon spin torque induced magnetization motion are presented and potential application examples are given. The structure and material of these proposed spin torque memristors are based upon existing (and/or commercialized) magnetic devices and can be easily integrated on top of a CMOS. This provides better controllability and flexibility to realize the promises of nanoscale memristors. Utilizing its unique device behavior, the paper explores spintronic memristor potential applications in multibit data storage and logic, novel sensing scheme, power management and information security.
Xiaobin Wang, Yiran Chen 0001
DATE1
2010 Variation tolerant sensing scheme of Spin-Transfer Torque Memory for yield improvement
abstract
Spin-Transfer Torque Random Access Memory (STT-RAM) demonstrated great potentials as an universal memory for its fast access speed, zero standby power, excellent scalability and simplicity of cell structure. However, large process variations of both magnetic tunneling junction and CMOS process severely limit the yield of STT-RAM chips and prevent the massive production from happening. In this paper, we analyze the impacts of process variations on various sensing schemes of STT-RAM. Based on our analysis, we propose a novel voltage-driven non-destructive self-reference sensing scheme (VDRS) to enhance the STT-RAM chip yield by significantly improving sense margin. Monte-Carlo simulations of a 16Kb STT-RAM array shows that VDRS can achieve the same yield as the previous non-destructive self-reference sensing scheme while improving the sense margin by 5.16 times with the similar access performance and power.
Zhenyu Sun 0001, Hai Li 0001, Yiran Chen 0001, Xiaobin Wang
ICCAD4
2010 Combined magnetic- and circuit-level enhancements for the nondestructive self-reference scheme of STT-RAM
abstract
A nondestructive self-reference read scheme (NSRS) was recently proposed to overcome the bit-to-bit variation in Spin-Transfer Torque Random Access Memory (STT-RAM). In this work, we introduced three magnetic- and circuit-level techniques, including 1) R-I curve skewing, 2) yield-driven sensing current selection, and 3) ratio matching to improve the sense margin and robustness of NSRS. The measurements of our 16Kb STT-RAM test chip show that compared to the original NSRS design, our proposed technologies successfully increased the sense margin by 2.5X with minimized impacts on the memory reliability and hardware cost.
Yiran Chen 0001, Hai Li 0001, Xiaobin Wang, Wenzhong Zhu, Wei Xu 0021, Tong Zhang 0002
ISLPED3
2010 Design Margin Exploration of Spin-Transfer Torque RAM (STT-RAM) in Scaled Technologies
abstract
We propose a magnetic and electric level spin-transfer torque random access memory (STT-RAM) cell model to simulate the write operation of an STT-RAM. The model of a magnetic tunneling junction (MTJ) is modified to take into account the electrical response of the MOS transistor that is connected to the MTJ. A dynamic design flow is also proposed to minimize any unnecessary design margin in an STT-RAM cell design by leveraging from the new STT-RAM cell model. The design of an STT-RAM cell with a one-transistor-one-MTJ (1T1J) structure shows that our technique can reduce more than 22% of the STT-RAM cell area, compared with a conventional STT-RAM cell model at a TSMC 90-nm technology node. The performance and the reliability of the memory cell were unaffected. By using our model, we analyzed the scalability of STT-RAM technology down to a 22-nm Bulk-CMOS technology node. The tradeoffs among the MTJ switching current, the thermal stability of the MTJ and the MOS transistor driving strength are discussed. Some magnetic- and circuit-level solutions to achieve 9F2STT-RAM cell area at 22-nm technology node are also discussed.
Yiran Chen 0001, Xiaobin Wang, Hai Li 0001, Haiwen Xi, Yuan Yan, Wenzhong Zhu
IEEE Trans. Very Large Scale Integr. Syst.2
2009 Improving STT MRAM storage density through smaller-than-worst-case transistor sizing
abstract
This paper presents a technique to improve the storage density of spin-torque transfer (STT) magnetoresistive random access memory (MRAM) in the presence of significant magnetic tunneling junction (MTJ) write current threshold variability. In conventional design practice, the nMOS transistor within each memory cell is sized to be large enough to carry a current larger than the worst-case MTJ write current threshold, leading to an increasing storage density penalty as the technology scales down. To mitigate such variability-induced storage density penalty, this paper presents a smaller-than-worst-case transistor sizing approach with the underlying theme of jointly considering memory cell transistor sizing and defect tolerance. Its effectiveness is demonstrated using 256Mb STT MRAM design at 45nm node as a test vehicle. Results show that, under a normalized write current threshold deviation of 20%, the overall memory die size can be reduced by more than 20% compared with the conventional worst-case transistor sizing design practice.
Wei Xu 0021, Yiran Chen 0001, Xiaobin Wang, Tong Zhang 0002
DAC3