VLDB 2026 Research / reviewers in the wild / expert
Shiwen Yu
dblp:11/652
· DBLP profile ↗
25ranked-venue papers
4as first author
4since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 1 first-authorSoftware engineering, systems software and programming languages · 3 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Neural Solving Uninterpreted Predicates with Abstract Gradient DescentabstractUninterpreted predicate solving is a fundamental problem in formal verification, including loop invariant and constrained horn clauses predicate solving. Existing approaches have been mostly in symbolic ways. While achieving sustainable progress, they still suffer from inefficiency and seem unable to leverage the ever-increasing computility, such as GPU. Recently, neural relaxation has been proposed to tackle this problem. They treat the uninterpreted predicate-solving task as an optimization problem by relaxing the discrete search process into a learning process of neural networks. However, two bottlenecks keep them from being valid. First, relaxed neural networks cannot match the original semantics of predicates rigorously; second, the neural networks are difficult to train to reach global optimization. Therefore, this article presents a novel discrete neural architecture with the Abstract Gradient Decent (AGD) algorithm to directly solve uninterpreted predicates in the discrete hypothesis space. The abstract gradient is for discrete neurons whose calculation rules are designed in an abstract domain. Our approach conforms to the original semantics of predicates, and the proposed AGD algorithm can achieve global optimization satisfactorily. We implement the tool Dasp in the Boxes abstract domain to solve uninterpreted predicates in the QF-NIA SMT theory. In the experiments, Dasp has outperformed seven state-of-the-art tools across three predicate synthesis tasks. Shiwen Yu, Zengyu Liu, Ting Wang 0009, Ji Wang 0001 |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2023 | Loop Invariant Inference through SMT Solving Enhanced Reinforcement LearningabstractInferring loop invariants is one of the most challenging problems in program verification. It is highly desired to incorporate machine learning when inferring. This paper presents a Reinforcement Learning (RL) pruning framework to infer loop invariants over a general nonlinear hypothesis space. The key idea is to synergize the RL-based pruning and SMT solving to generate candidate invariants efficiently. To address the sparse reward problem in learning, we design a novel two-dimensional reward mechanism that enables the RL pruner to recognize the capability boundary of SMT solvers and learn the pruning heuristics in a few rounds. We have implemented our approach with Z3 SMT solver in the tool called LIPuS and conducted extensive experiments over the linear and nonlinear benchmarks. Experiment results show that LIPuS can solve the most cases compared to the state-of-the-art loop invariant inference tools such as Code2Inv, ICE-DT, GSpacer, SymInfer, ImplCheck, and Eldarica. Especially, LIPuS outperforms them significantly on nonlinear benchmarks. Shiwen Yu, Ting Wang 0009, Ji Wang 0001 |
ISSTA | 1 |
| 2022 | HeHe: Balancing the Privacy and Efficiency in Training CNNs over the Semi-honest Cloud
Longlong Sun, Hui Li 0006, Shiwen Yu, XinDi Ma, Yanguo Peng, Jiangtao Cui |
ISC | 3 |
| 2022 | Data Augmentation by Program Transformation
Shiwen Yu, Ting Wang 0009, Ji Wang 0001 |
J. Syst. Softw. | 1 |
| 2019 | Multi-classification of Theses to Disciplines Based on Metadata
Jianling Li, Shiwen Yu, Shasha Li 0001, Jie Yu 0008 |
NLPCC (2) | 2 |
| 2012 | Societally connected multimedia across culturesabstractThe advance of the Internet in the past decade has radically changed the way people communicate and collaborate with each other. Physical distance is no more a barrier in online social networks, but cultural differences (at the individual, community, as well as societal levels) still govern human-human interactions and must be considered and leveraged in the online world. The rapid deployment of high-speed Internet allows humans to interact using a rich set of multimedia data such as texts, pictures, and videos. This position paper proposes to define a new research area called ‘connected multimedia’, which is the study of a collection of research issues of the super-area social media that receive little attention in the literature. By connected multimedia, we mean the study of the social and technical interactions among users, multimedia data, and devices across cultures and explicitly exploiting the cultural differences. We justify why it is necessary to bring attention to this new research area and what benefits of this new research area may bring to the broader scientific research community and the humanity. Zhongfei Zhang, Zhengyou Zhang, Ramesh Jain 0001, Yueting Zhuang, Noshir S. Contractor, Alex Hauptmann 0001, Alejandro Jaimes, Wanqing Li 0001, Alexander C. Loui, Tao Mei 0001, Nicu Sebe, Yonghong Tian 0001, Vincent S. Tseng, Qing Wang 0015, Changsheng Xu, Shiwen Yu |
J. Zhejiang Univ. Sci. C | 17 |
| 2010 | Automatic Acquisition of Chinese Novel Noun Compounds
Chu-Ren Huang, Shiwen Yu |
LREC | 3 |
| 2008 | Text normalization in mandarin text-to-speech systemabstractText normalization is an important component in text-to-speech system and the difficulty in text normalization is to disambiguate the non-standard words (NSWs). This paper develops a taxonomy of NSWs on the basis of a large scale Chinese corpus, and proposes a two-stage NSWs disambiguation strategy, finite state automata (FSA) for initial classification and maximum entropy (ME) classifiers for subclass disambiguation. Based on the above NSWs taxonomy, the two-stage approach achieves an F-score of 98.53% in open test, 5.23% higher than that of FSA based approach. Experiments show that the NSWs taxonomy ensures FSA a high baseline performance and ME classifiers make considerable improvement, and the two-stage approach adapts well to new domains. Yuxiang Jia, Dezhi Huang, Shiwen Yu, Haila Wang |
ICASSP | 4 |
| 2008 | Quality Assurance of Automatic Annotation of Very Large Corpora: a Study based on heterogeneous Tagging System
Chu-Ren Huang, Lung-Hao Lee, Jia-Fei Hong, Weiguang Qu, Shiwen Yu |
LREC | 5 |
| 2008 | Unsupervised Chinese Verb Metaphor Recognition Based on Selectional Preferences
Yuxiang Jia, Shiwen Yu |
PACLIC | 2 |
| 2007 | Word Clustering for Collocation-Based Word Sense Disambiguation
Xu Sun 0001, Yunfang Wu, Shiwen Yu |
CICLing | 4 |
| 2007 | A Collocation-Based WSD Model: RFR-SUM
Weiguang Qu, Zhifang Sui, Genlin Ji, Shiwen Yu, Junsheng Zhou |
IEA/AIE | 4 |
| 2006 | Chinese Noun Phrase Metaphor Recognition with Maximum Entropy Approach
Houfeng Wang, Huiming Duan, Shiwen Yu |
CICLing | 5 |
| 2006 | A Comparative Study on Representing Units in Chinese Text Clustering
Shiwen Yu, Xueqiang Lv, Shuicai Shi, Shibin Xiao |
KSEM | 2 |
| 2005 | Extracting Terminologically Relevant Collocations in the Translation of Chinese Monograph
Byeong Kwu Kang, Baobao Chang, Yi-Rong Chen, Shiwen Yu |
IJCNLP | 4 |
| 2004 | A Combining Approach to Automatic Keyphrases Indexing for Chinese News Documents
Houfeng Wang, Sujian Li, Shiwen Yu, Byeong Kwu Kang |
CICLing | 3 |
| 2004 | Distributional Consistency: As a General Method for Defining a Core Lexicon
Huarui Zhang, Chu-Ren Huang, Shiwen Yu |
LREC | 3 |
| 2004 | An adaptive k-nearest neighbor text categorization strategyabstractk is the most important parameter in a text categorization system based on the k -nearest neighbor algorithm ( k NN). To classify a new document, the k -nearest documents in the training set are determined first. The prediction of categories for this document can then be made according to the category distribution among the k nearest neighbors. Generally speaking, the class distribution in a training set is not even; some classes may have more samples than others. The system's performance is very sensitive to the choice of the parameter k . And it is very likely that a fixed k value will result in a bias for large categories, and will not make full use of the information in the training set. To deal with these problems, an improved kNN strategy, in which different numbers of nearest neighbors for different categories are used instead of a fixed number across all categories, is proposed in this article. More samples (nearest neighbors) will be used to decide whether a test document should be classified in a category that has more samples in the training set. The numbers of nearest neighbors selected for different categories are adaptive to their sample size in the training set. Experiments on two different datasets show that our methods are less sensitive to the parameter k than the traditional ones, and can properly classify documents belonging to smaller classes with a large k . The strategy is especially applicable and promising for cases where estimating the parameter k via cross-validation is not possible and the class distribution of a training set is skewed. Baoli Li 0001, Qin Lu 0001, Shiwen Yu |
ACM Trans. Asian Lang. Inf. Process. | 3 |
| 2003 | Experimental Study on Representing Units in Chinese Text Categorization
Baoli Li 0001, Xiaojing Bai, Shiwen Yu |
CICLing | 4 |
| 2003 | News-Oriented Keyword Indexing with Maximum Entropy Principle
Sujian Li, Houfeng Wang, Shiwen Yu, Chengsheng Xin |
PACLIC | 3 |
| 2003 | A Large-scale Lexical Semantic Knowledge-base of Chinese
Hui Wang 0009, Shiwen Yu |
PACLIC | 2 |
| 2002 | Building a Bilingual WordNet-Like Lexicon: The New Approach and Algorithms
Shiwen Yu, Jiangsheng Yu |
COLING | 2 |
| 2000 | The Multi-layer Language Knowledge Base of Chinese NLP
Shiwen Yu |
LREC | 2 |
| 1994 | Blending Segmentation With Tagging In Chinese Language Corpus Processing
Shiwen Yu |
COLING | 2 |
| 1993 | Automatic evaluation of output quality for Machine Translation systems
Shiwen Yu |
Mach. Transl. | 1 |