Wensheng Wu

dblp:23/7 · DBLP profile ↗
← Back
30ranked-venue papers
18as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 12 · 6 first-authorHuman-computer interaction and ubiquitous computing · 10 · 10 first-author · 9 since 2021Artificial intelligence and machine learning · 9 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Software engineering, systems software and programming languages · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Estimating total organic carbon by a graph convolutional prediction model considering geological context weight
Wensheng Wu, Zhangxin Chen, Benjieming Liu
Eng. Appl. Artif. Intell.2
2026 Improving anomaly detection with foundation-model synthesis and wavelet-domain attention
Wensheng Wu, Zheming Lu 0001, Ziqian Lu, Zewei He, Xuecheng Sun, Jungong Han, Yunlong Yu 0001
Neural Networks1
2024 Learning Big Data Systems via Emulation
abstract
Big data systems are becoming an integral part of computing and data science curriculum. However, the current curriculum is largely focused on how to use the systems. An effective approach to learning the internals of big data systems is through emulation. In this paper, we report on a study where students in a graduate database course were asked to complete a course project on emulating big data systems such as Hadoop and Spark. We present the design of the emulation projects and examine the impact of the projects on students' learning. Our key finding is that the emulation projects can greatly improve students' self-efficacy in completing tasks that require in-depth knowledge and skills on big data systems.
Wensheng Wu
SIGCSE (1)1
2024 Well Logs Reconstruction Based on Deep Learning Technology
abstract
Well logs play a very important role in formation evaluation. But due to the influence of geological, engineering and other factors in practical applications, some logs are often distorted or missing. The log curve reconstruction methods based on traditional empirical model and statistical analysis are unconvincing enough, so it is proposed to reconstruct the log curves by deep learning method. Considering the limitation of neural network, LSTM (long short-term memory) neural network is studied. And we utilize the formation lithology index (FLI) of logging domain knowledge to filter data. As a result, high-quality training sample data are acquired, which are used as the basis for deep learning to reconstruct log curves. Meanwhile, the LSTM model with domain knowledge constraint layer is constructed and trained. Then PSO algorithm is chosen to optimize the LSTM hyper-parameters—initial learning rate and dropout probability. Finally, a log curve reconstruction model (PSO-DK-LSTM) combining domain knowledge and PSO algorithm is created. We design one experiment to verify prediction capability of PSO-DK-LSTM. The whole data are collected from a work area of XX oilfield. The verification results show that the log reconstruction method based on PSO-DK-LSTM has achieved better results in terms of accuracy and stability, which gives a new idea for log reconstruction.
Wensheng Wu
IEEE Geosci. Remote. Sens. Lett.2
2023 Towards a Validated Self-Efficacy Scale for Data Management
abstract
We propose a self-efficacy scale for data management. The scale assesses students' perceived capabilities in mastering the breadth and depth of modern data management, as well as hands-on skills for effective management of data. Such capabilities are critical to computing and data science students. We have conducted factor analysis to validate the scale. The analysis produced a factor model with high internal consistencies. Group analyses using the factor solution and statistical testing show that (1) males and females have similar self-efficacy, except for the depth of knowledge where females showed higher confidences; and (2) CS students had much higher self-efficacy than non-CS students. To the best of our knowledge, this is the first self-efficacy scale for data management.
Wensheng Wu
SIGCSE (1)1
2023 Assessing Peer Correction of SQL and NoSQL Queries
abstract
Students in database courses often make varied syntax and semantic mistakes in writing SQL and NoSQL queries. We report on a study where we designed two styles of multiple-choice questions based on students' mistakes in midterms and used the questions in the final exam to assess student's capabilities in identifying and correcting other students' mistakes. The study found that (1) students had similar performance on both styles of the questions; (2) the average accuracy rate of students in peer correction was about 83%; (3) over 80% of the students thought it was helpful to see and correct others' mistakes; and (4) students' performance in peer correction was moderately correlated with their overall performance. This study is the first to address students' mistakes in writing NoSQL queries and assess peer correction of both SQL and NoSQL queries.
Wensheng Wu
SIGCSE (1)1
2022 Data Science Course Projects with Peer Challenges: An Experience Report
abstract
Project-based learning is a key part of learning experiences for data science students. Typically, students choose their own topics and work on the projects without much interaction among the different groups. In this paper, we present a new approach to project design where every group is asked to also create a data science challenge based on its project idea, and these peer challenges are assigned to multi-person groups to tackle as additional requirements for their projects. This approach provides new peer learning opportunities to the students through challenge formulation, solving, and assessment. We have applied the approach to design course projects for a database course in our data science program with students from diverse academic backgrounds. Results from peer evaluations and surveys indicate that students were generally satisfied with the quality of peer challenges and solutions, and the process of formulating and solving challenges has helped students gain problem design, critical thinking, and data analytics skills.
Wensheng Wu
ITiCSE (1)1
2022 Investigating Internship Experiences of Data Science Students for Curriculum Enhancement
abstract
An internship promotes the idea of learning through experiencing and has been an important way for data science students to gain practical knowledge and skills for solving real-world data science challenges. However, despite its importance, there has been little work on understanding student internship experiences. To address this knowledge gap, we have conducted a study to find out how the current curriculum has prepared students for their internships, what are the major challenges students have faced in their internships, and how to enhance the curriculum to better prepare students for their internships and jobs. In this paper, we report findings of our study and discuss strategies for improving the curriculum based on student experiences and feedback. Our study provides key insights on student learning experiences that may benefit other data science programs. Our work also constitutes an important step towards establishing a feedback loop that couples internships and jobs with the curriculum, and continuously adapts the curriculum to the fast-evolving techniques and tools used in data science applications.
Wensheng Wu
ITiCSE (1)1
2021 One Size Doesn't Fit All: Diversifying Data Science Course Projects by Student Background and Interests
abstract
A key challenge to a data science program is how to tailor its curriculum to accommodate the needs and interests of students from diverse academic backgrounds. To address this challenge, we propose Proj4X, a general project model that gives students the flexibility in choosing subject areas for the projects based on their background and interests, while requiring every project to address the key knowledge and skills covered in the course. Although project-based learning has been well researched, only a few studies have addressed the challenges in managing student-initiated projects. We have applied Proj4X to a graduate database course in our data science program. We used student peer evaluation to measure the performance of projects objectively and study the effect of group size on the performance of projects and the correlation between specific aspects of a project and its overall rating. Student feedback shows that Proj4X can greatly foster the collaboration among students of diverse backgrounds and inspire them to work on a wide range of interesting topics: from satellite tracker, hospital search during pandemic, to wildlife diversity monitoring.
Wensheng Wu
ITiCSE (1)1
2021 CS vs non-CS: Analyzing Online Social Behaviors of Data Science Students with Diverse Academic Backgrounds
abstract
Understanding the characteristics of diverse students in a data science program is a key to the success of the program. Towards this goal, we have conducted a study to understand social behaviors of data science students on online platforms, in particular, Piazza and Zoom. Our key findings are: (1) Math/statistics students were least active on Zoom chats, but frequent questioners on Piazza, which suggests that we need to engage these students more during Zoom meetings. (2) Most of endorsed answers on Piazza were contributed by CS students, which suggests a collaborative culture formed between CS and non-CS students. (3) The number of Zoom chats and activities on Piazza correlated positively with student performance, but much more strongly for non-CS students than CS-students.
Wensheng Wu
ITiCSE (2)1
2021 SQL2X: Learning SQL, NoSQL, and MapReduce via Translation
abstract
A key challenge in designing a database course is how to introduce students to the great variety of data models, query languages, databases, and data processing systems available now. To address this challenge, we propose SQL2X, a novel SQL-centric learning model that teaches students SQL, NoSQL, and MapReduce via translation. For example, translating SQL queries into MapReduce programs to gain insights on how aggregation and join are performed in parallel in MapReduce, and translating SQL queries into REST requests to Firebase to help understand the differences between the query capability of SQL and NoSQL databases. We have applied the model to a graduate database course in our applied data science program. The evaluation and feedback from the students with diverse background indicate the effectiveness of the model in developing students' modeling, querying, and analytical skills over diverse data systems.
Wensheng Wu
SIGCSE1
2021 Proj4X: Diversifying Data Science Course Projects by Student Background and Interests
abstract
A key challenge to a data science program is how to tailor its curriculum to accommodate the needs and interests of students with diverse backgrounds. To address this challenge, we propose Proj4X, a model for diversifying the topics and structures of course projects based on student background and interests. Our study shows that the model can greatly foster the collaboration among the students and motivate them in acquiring the necessary knowledge and skills for solving a great variety of real-world problems ranging from COVID-19, CO2 emission, to wildlife diversity.
Wensheng Wu
SIGCSE1
2018 Software Architecture Measurement - Experiences from a Multinational Company
Wensheng Wu, Yuanfang Cai, Rick Kazman, Ran Mo, Rongbiao Chen, Yingan Ge, Weicai Liu
ECSA1
2018 SLASH: Automatically Generating Flash Cards for Reviewing Concepts in Lectures Slides (Abstract Only)
abstract
We present SLASH, a learning tool currently under development in our graduate program. SLASH aims to help students review concepts in lectures slides using flash cards automatically generated from the slides. Many courses in our program have weekly quizzes and students can get stressed quite easily. So we hope that SLASH can make the process of reviewing lectures more fun and interesting to the students. Extracting concepts from lectures slides is itself an interesting but challenging problem, since the contents of the slides may be fragmented (e.g., point-based, with an incomplete sentence for each point) and noisy (e.g., containing formulas and codes). Past research on text mining has tried to "glue" together the points to construct a grammatically correct sentence, which is then used to extract concepts and relationships. In contrast, we focus on discovering popular concepts in the slides and generating flash cards with (just) sufficient contexts to help students recall the concepts. To the best of our knowledge, this is the first work on the automatic generation of concept-based flash cards from lecture slides. In the presentation, we will show our preliminary work, example flash cards, student feedback, and challenges in developing SLASH. We believe that SLASH may benefit all instructors who are using PowerPoint for lecture presentation, and may be used to largely stimulate students' interests in learning the subjects.
Wensheng Wu
SIGCSE1
2016 Learning semantic representation with neural networks for community question answering retrieval
Guangyou Zhou, Tingting He 0003, Wensheng Wu
Knowl. Based Syst.4
2016 Q2P: Discovering Query Templates via Autocompletion
abstract
We present Q2P, a system that discovers query templates from search engines via their query autocompletion services. Q2P is distinct from the existing works in that it does not rely on query logs of search engines that are typically not readily available. Q2P is also unique in that it uses a trie to economically store queries sampled from a search engine and employs a beam-search strategy that focuses the expansion of the trie on its most promising nodes. Furthermore, Q2P leverages the trie-based storage of query sample to discover query templates using only two passes over the trie. Q2P is a key part of our ongoing project Deep2Q on a template-driven data integration on the Deep Web, where the templates learned by Q2P are used to guide the integration process in Deep2Q. Experimental results on four major search engines indicate that (1) Q2P sends only a moderate number of queries (ranging from 597 to 1,135) to the engines, while obtaining a significant number of completions per query (ranging from 4.2 to 8.5 on the average); (2) a significant number of templates (ranging from 8 to 32 when the minimum support for frequent templates is set to 1%) may be discovered from the samples.
Wensheng Wu, Weiyi Meng, Weifeng Su, Guangyou Zhou, Yao-Yi Chiang
ACM Trans. Web1
2015 A Subspace Learning Framework for Cross-Lingual Sentiment Classification with Partial Parallel Data
Guangyou Zhou, Tingting He 0003, Jun Zhao 0001, Wensheng Wu
IJCAI4
2015 Linking Heterogeneous Input Features with Pivots for Domain Adaptation
Guangyou Zhou, Tingting He 0003, Wensheng Wu, Xiaohua Hu 0001
IJCAI3
2014 An empirical study of topic-sensitive probabilistic model for expert finding in question answer communities
Guangyou Zhou, Jun Zhao 0001, Tingting He 0003, Wensheng Wu
Knowl. Based Syst.4
2013 Semantic discovery from web comparison queries
abstract
Users frequently pose comparison queries (e.g., ibm vs apple) on web search engines. However, little research has been done on understanding these queries. To fill in this gap, this paper describes a first solution to discovering and mining comparison queries. We present a novel snowballing algorithm that "crawls" comparison queries from search engines via their query autocompletion services. We propose a novel modeling approach that represents comparison queries in a comparison graph and develop a novel algorithm that mines closely related concepts from comparison graphs via spectral clustering. Initial experiments indicate that our approach can reveal the inherent semantic relationship among the concepts and discover different senses of a concept, e.g., "toyota" as a car brand or a company name.
Tingting Zhong, Wensheng Wu
CIKM2
2013 Proactive natural language search engine: tapping into structured data on the web
abstract
In this era of "big data", a key challenge facing the database community is to help average users tap into the huge amounts of structured data on the Web. To address this challenge, we propose a novel proactive template-based engine for searching structured data on the Web using natural language. Departing from conventional search engines, the proposed engine organizes questions it can answer using templates and figures out ahead of time which sources can answer which templates and how. Then, at query time, the engine can simply match queries with the templates and retrieve answers using the pre-compiled evaluation plans. While attractive, building such an engine requires innovations in template creation, query evaluation, and system evolution. In this paper, we propose novel techniques to address these challenges.
Wensheng Wu
EDBT1
2008 Discovering topical structures of databases
abstract
The increasing complexity of enterprise databases and the prevalent lack of documentation incur significant cost in both understanding and integrating the databases. Existing solutions addressed mining for keys and foreign keys, but paid little attention to more high-level structures of databases. In this paper, we consider the problem of discovering topical structures of databases to support semantic browsing and large-scale data integration. We describe iDisc, a novel discovery system based on a multi-strategy learning framework. iDisc exploits varied evidence in database schema and instance values to construct multiple kinds of database representations. It employs a set of base clusterers to discover preliminary topical clusters of tables from database representations, and then aggregate them into final clusters via meta-clustering. To further improve the accuracy, we extend iDisc with novel multiple-level aggregation and clusterer boosting techniques. We introduce a new measure on table importance and propose an approach to discovering cluster representatives to facilitate semantic browsing. An important feature of our framework is that it is highly extensible, where additional database representations and base clusterers may be easily incorporated into the framework. We have extensively evaluated iDisc using large real-world databases and results show that it discovers topical structures with a high degree of accuracy.
Wensheng Wu, Berthold Reinwald, Yannis Sismanis, Rajesh Manjrekar
SIGMOD Conference1
2006 Merging Source Query Interfaces onWeb Databases
abstract
Recently, there are many e-commerce search engines that return information from Web databases. Unlike text search engines, these e-commerce search engines have more complicated user interfaces. Our aim is to construct automatically a natural query user interface that integrates a set of interfaces over a given domain of interest. For example, each airline company has a query interface for ticket reservation and our system can construct an integrated interface for all these companies. This will permit users to access information uniformly from multiple sources. Each query interface from an e-commerce search engine is designed so as to facilitate users to provide necessary information. Specifically, (1) related pieces of information such as first name and last name are grouped together and (2) certain hierarchical relationships are maintained. In this paper, we provide an algorithm to compute an integrated interface from query interfaces of the same domain. The integrated query interface can be proved to preserve the above two types of relationships. Experiments on five domains verify our theoretical study.
Eduard C. Dragut, Wensheng Wu, A. Prasad Sistla, Clement T. Yu, Weiyi Meng
ICDE2
2006 WebIQ: Learning from the Web to Match Deep-Web Query Interfaces
abstract
Integrating Deep Web sources requires highly accurate semantic matches between the attributes of the source query interfaces. These matches are usually established by comparing the similarities of the attributes’ labels and instances. However, attributes on query interfaces often have no or very few data instances. The pervasive lack of instances seriously reduces the accuracy of current matching techniques. To address this problem, we describe WebIQ, a solution that learns from both the Surface Web and the Deep Web to automatically discover instances for interface attributes. WebIQ extends question answering techniques commonly used in the AI community for this purpose. We describe how to incorporate WebIQ into current interface matching systems. Extensive experiments over five realworld domains show the utility ofWebIQ. In particular, the results show that acquired instances help improve matching accuracy from 89.5% F-1 to 97.5%, at only a modest runtime overhead.
Wensheng Wu, AnHai Doan, Clement T. Yu
ICDE1
2005 Merging Interface Schemas on the Deep Web via Clustering Aggregation
abstract
We consider the problem of integrating a large number of interface schemas over the deep Web, The scale of the problem and the diversity of the sources present serious challenges to the conventional manual or rule-based approaches to schema integration. To address these challenges, we propose a novel formulation of schema integration as an optimization problem, with the objective of maximally satisfying the constraints given by individual schemas. Since the optimization problem can be shown to be NP-complete, we develop a novel approximation algorithm LMax, which builds the unified schema via recursive applications of clustering aggregation. We further extend LMax to handle the irregularities frequently occurring among the interface schemas. Extensive evaluation on real-world data sets shows the effectiveness of our approach.
Wensheng Wu, AnHai Doan, Clement T. Yu
ICDM1
2004 An Interactive Clustering-based Approach to Integrating Source Query interfaces on the Deep Web
abstract
An increasing number of data sources now become available on the Web, but often their contents are only accessible through query interfaces. For a domain of interest, there often exist many such sources with varied coverage or querying capabilities. As an important step to the integration of these sources, we consider the integration of their query interfaces. More specifically, we focus on the crucial step of the integration: accurately matching the interfaces. While the integration of query interfaces has received more attentions recently, current approaches are not sufficiently general: (a) they all model interfaces with flat schemas; (b) most of them only consider 1:1 mappings of fields over the interfaces; (c) they all perform the integration in a blackbox-like fashion and the whole process has to be restarted from scratch if anything goes wrong; and (d) they often require laborious parameter tuning. In this paper, we propose an interactive, clustering-based approach to matching query interfaces. The hierarchical nature of interfaces is captured with ordered trees. Varied types of complex mappings of fields are examined and several approaches are proposed to effectively identify these mappings. We put the human integrator back in the loop and propose several novel approaches to the interactive learning of parameters and the resolution of uncertain mappings. Extensive experiments are conducted and results show that our approach is highly effective.
Wensheng Wu, Clement T. Yu, AnHai Doan, Weiyi Meng
SIGMOD Conference1
2002 A Statistical Method for Estimating the Usefulness of Text Databases
abstract
Searching desired data on the Internet is one of the most common ways the Internet is used. No single search engine is capable of searching all data on the Internet. The approach that provides an interface for invoking multiple search engines for each user query has the potential to satisfy more users. When the number of search engines under the interface is large, invoking all search engines for each query is often not cost effective because it creates unnecessary network traffic by sending the query to a large number of useless search engines and searching these useless search engines wastes local resources. The problem can be overcome if the usefulness of every search engine with respect to each query can be predicted. We present a statistical method to estimate the usefulness of a search engine for any given query. For a given query, the usefulness of a search engine in this paper is defined to be a combination of the number of documents in the search engine that are sufficiently similar to the query and the average similarity of these documents. Experimental results indicate that our estimation method is much more accurate than existing methods.
King-Lup Liu, Clement T. Yu, Weiyi Meng, Wensheng Wu, Naphtali Rishe
IEEE Trans. Knowl. Data Eng.4
2001 Efficient and Effective Metasearch for Text Databases Incorporating Linkages among Documents
abstract
Linkages among documents have a significant impact on the importance of documents, as it can be argued that important documents are pointed to by many documents or by other important documents. Metasearch engines can be used to facilitate ordinary users for retrieving information from multiple local sources (text databases). There is a search engine associated with each database. In a large-scale metasearch engine, the contents of each local database is represented by a representative. Each user query is evaluated against he set of representatives of all databases in order to determine the appropriate databases (search engines) to search (invoke) In previous word, the linkage information between documents has not been utilized in determining the appropriate databases to search. In this paper, such information is employed to determine the degree of relevance of a document with respect to a given query. Specifically, the importance (rank) of each document as determined by the linkages is integrated in each database representative to facilitate the selection of databases for each given query. We establish a necessary and sufficient condition to rank databases optimally, while incorporating the linkage information. A method is provided to estimate the desired quantities stated in the necessary and sufficient condition. The estimation method runs in time linearly proportional to the number of query terms. Experimental results are provided to demonstrate the high retrieval effectiveness of the method.
Clement T. Yu, Weiyi Meng, Wensheng Wu, King-Lup Liu
SIGMOD Conference3
1999 Efficient and Effective Metasearch for a Large Number of Text Databases
abstract
Metasearch engines can be used to facilitate ordinary users for retrieving information from multiple local sources (text databases). In a metasearch engine, the contents of each local database is represented by a representative. Each user query is evaluated against the set of representatives of all databases in order to determine the appropriate databases to search. When the number of databases is very large, say in the order of tens of thousands or more, then a traditional metasearch engine may become inefficient as each query needs to be evaluated against too many database representatives. Furthermore, the storage requirement on the site containing the metasearch engine can be very large. In this paper, we propose to use a hierarchy of database representatives to improve the efficiency. We provide an algorithm to search the hierarchy. We show that the retrieval effectiveness of our algorithm is the same as that of evaluating the user query against all database representatives. We also show that our algorithm is efficient. In addition, we propose an alternative way of allocating representatives to sites so that the storage burden on the site containing the metasearch engine is much reduced.
Clement T. Yu, Weiyi Meng, King-Lup Liu, Wensheng Wu, Naphtali Rishe
CIKM4
1999 Estimating the Usefulness of Search Engines
abstract
In this paper, we present a statistical method to estimate the usefulness of a search engine for any given query. The estimates can be used by a metasearch engine to choose local search engines to invoke. For a given query, the usefulness of a search engine in this paper is defined to be a combination of the number of documents in the search engine that are sufficiently similar to the query and the average similarity of these documents. Experimental results indicate that the proposed estimation method is quite accurate.
Weiyi Meng, King-Lup Liu, Clement T. Yu, Wensheng Wu, Naphtali Rishe
ICDE4