VLDB 2026 Research / reviewers in the wild / expert
Shuguang Wang
dblp:22/2874
· DBLP profile ↗
28ranked-venue papers
8as first author
14since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 9 · 3 first-author · 2 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 4 · 3 since 2021Computer networks · 2 · 2 since 2021Security and privacy · 2 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Developing Authentic Simulated Learners for Mathematics Teacher Learning: Insights from Three Approaches with Large Language Models
Selim Yavuz, Boran Yu, Shuguang Wang, Pavneet Kaur Bharaj, Dionne Cross Francis |
AIED (5) | 5 |
| 2026 | Latent diffusion-driven inverse design of damping microstructures with multiaxial nonlinear mechanical targets
Tianyang Zhang 0013, Weizhi Xu 0005, Shuguang Wang, Dongsheng Du |
Adv. Eng. Informatics | 3 |
| 2026 | When, Who, and Why: Exploring Occupants' Demand of Explanations from Autonomous VehiclesabstractWhile autonomous vehicles (AVs) could transform transportation, their “black box” nature often leaves occupants unaware of the rationale for their actions. Providing explanations can enhance transparency and facilitate widespread AV adoption. This study investigated scenario (when) and human factors (who) that influence the Demand of Explanations (DoE) to ensure explanations are provided when needed, followed by exploring the reasons (why) behind these demands. We conducted an online experimental study among 440 participants, who viewed 36 simulated driving scenarios, varying in AV actions, driving styles, time/weather and traffic environments. Results of multilevel and qualitative analysis showed that: (1) DoE was significantly higher when AVs drove aggressively, in urban areas, during turning and merging, and in nighttime or rain; (2) participants who had lower trust in AVs and older adults significantly demanded more explanations; and (3) safety and traffic rules were the primary reasons for seeking explanations. Yilin Kou, Qian Zhou 0008, Shuguang Wang, Nancy Xiaonan Yu, Zhicong Lu, Jianping Wang 0001 |
Int. J. Hum. Comput. Interact. | 3 |
| 2026 | Opportunistic Gossip Learning-Aided Collaborative Physical Layer Authentication for Internet of VehiclesabstractTraditional device authentication techniques relying only on a single device for authentication are susceptible to poor performance due to insufficient channel observations. Although these challenges are mitigated by centralized collaborative authentication schemes, such schemes assume that all users desire to learn the same model. They further assume constant collaborator presence, leading to poorly trained models in Internet of Vehicles (IoV) environments with fleeting encounters. Moreover, they are vulnerable to GPS spoofing attack. To overcome these limitations, we propose a Gossip Learning (GL)-based collaborative Physical Layer Authentication (PLA) scheme, where collaborators assist vehicles in authenticating transmitters by distributively training their Long Short-term Memory (LSTM) models using the location and Channel State Information (CSI) of the transmitters. The collaborators also improve their model instances by incorporating the learned experiences of other vehicles they meet opportunistically. Instead of using the claimed location of transmitters during authentication, which may be affected by the GPS spoofing attack, their Received Signal Strength (RSS) and Angle of Arrival (AoA) features are used to estimate their current location. The estimated location is then used to predict their current CSI for authentication. Moreover, we propose a distance-based GPS signal spoofing detection algorithm, which ensures that only collaborators not under GPS spoofing attack are selected for collaborative training. The simulation results from experiments conducted using realistic channel attributes obtained from the Quasi-Deterministic Radio channel Generator (QuaDRiGa) platform demonstrate the effectiveness of our system in IoV scenarios, underscoring its relevance to Intelligent Transportation Systems and its superiority over existing techniques. Mubarak Umar, Shuangrui Zhao, Shuguang Wang, Ming-Gang Zheng, Yulong Shen 0001, Zhiwei Zhang 0004, Yufeng Kang, Hafsa Kabir Ahmad |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2025 | Interventional Root Cause Analysis of Failures in Multi-Sensor Fusion Perception Systems
Shuguang Wang, Qian Zhou 0008, Kui Wu 0001, Jinghuai Deng, Dapeng Oliver Wu, Wei-Bin Lee, Jianping Wang 0001 |
NDSS | 1 |
| 2025 | Efficient Pre-Training of LLMs via Topology-Aware Communication Alignment on More Than 9600 GPUsabstractThe scaling law for large language models (LLMs) depicts that the path towards machine intelligence necessitates training at large scale. Thus, companies continuously build large-scale GPU clusters, and launch training jobs that span over thousands of computing nodes. However, LLM pre-training presents unique challenges due to its complex communication patterns, where GPUs exchange data in sparse yet high-volume bursts within specific groups. Inefficient resource scheduling exacerbates bandwidth contention, leading to suboptimal training performance. This paper presents Arnold, a scheduling system summarizing our experience to effectively align LLM communication patterns to data center topology at scale. In-depth characteristic study is performed to identify the impact of physical network topology to LLM pre-training jobs. Based on the insights, we develop a scheduling algorithm to effectively align communication patterns to physical network topology in data centers. Through simulation experiments, we show the effectiveness of our algorithm in reducing the maximum spread of communication groups by up to $1.67$x. In production training, our scheduling system improves the end-to-end performance by $10.6\%$ when training with more than $9600$ Hopper GPUs, a significant improvement for our training pipeline. Youhe Jiang, Wencong Xiao, Kaihua Jiang, Shuguang Wang, Jun Wang 0039, Zixian Du, Zhuo Jiang, Binhang Yuan, Eiko Yoneki |
NeurIPS | 5 |
| 2025 | REDOUBT: Duo Safety Validation for Autonomous Vehicle Motion PlanningabstractSafety validation, which assesses the safety of an autonomous system's motion planning decisions, is critical for the safe deployment of autonomous vehicles. Existing input validation techniques from other machine learning domains, such as image classification, face unique challenges in motion planning due to its contextual properties, including complex inputs and one-to-many mapping. Furthermore, current output validation methods in autonomous driving primarily focus on open-loop trajectory prediction, which is ill-suited for the closed-loop nature of motion planning. We introduce REDOUBT, the first systematic safety validation framework for autonomous vehicle motion planning that employs a duo mechanism, simultaneously inspecting input distributions and output uncertainty. REDOUBT identifies previously overlooked unsafe modes arising from the interplay of In-Distribution/Out-of-Distribution (OOD) scenarios and certain/uncertain planning decisions. We develop specialized solutions for both OOD detection via latent flow matching and decision uncertainty estimation via an energy-based approach. Our extensive experiments demonstrate that both modules outperform existing approaches, under both open-loop and closed-loop evaluation settings. Our codes are available at: https://github.com/sgNicola/Redoubt. Shuguang Wang, Qian Zhou 0008, Kui Wu 0001, Dapeng Oliver Wu, Wei-Bin Lee, Jianping Wang 0001 |
NeurIPS | 1 |
| 2025 | Minder: Faulty Machine Detection for Large-scale Distributed Model Training
Yangtao Deng, Zhuo Jiang, Xingjian Zhang 0009, Zhang Zhang 0003, Zuquan Song, Gaohong Liu, Fuliang Li, Shuguang Wang, Haibin Lin, Jianxi Ye, Minlan Yu |
NSDI | 12 |
| 2025 | Understanding Stragglers in Large Model Training Using What-if Analysis
Jinkun Lin, Ziheng Jiang, Zuquan Song, Sida Zhao, Menghan Yu, Zhanghan Wang, Zuocheng Shi, Zherui Liu, Shuguang Wang, Haibin Lin, Xin Liu 0086, Aurojit Panda, Jinyang Li 0001 |
OSDI | 12 |
| 2025 | Mycroft: Tracing Dependencies in Collective Communication Towards Reliable LLM TrainingabstractReliability is essential for ensuring efficiency in LLM training. However, many real-world reliability issues remain difficult to resolve, resulting in wasted resources and degraded model performance. Unfortunately, today's collective communication libraries operate as black boxes, hiding critical information needed for effective root cause analysis. Yangtao Deng, Qinlong Wang, Xiaoyun Zhi, Zhuo Jiang, Haohan Xu, Zuquan Song, Gaohong Liu, Shuguang Wang, Wencong Xiao, Jianxi Ye, Minlan Yu, Hong Xu 0001 |
SOSP | 12 |
| 2025 | Robust LLM Training Infrastructure at ByteDanceabstractThe training scale of large language models (LLMs) has reached tens of thousands of GPUs and is still continuously expanding, enabling faster learning of larger models. Accompanying the expansion of the resource scale is the prevalence of failures (CUDA error, NaN values, job hang, etc.), which poses significant challenges to training stability. Any large-scale LLM training infrastructure should strive for minimal training interruption, efficient fault diagnosis, and effective failure tolerance to enable highly efficient continuous training. This paper presents ByteRobust, a large-scale GPU infrastructure management system tailored for robust and stable training of LLMs. It exploits the uniqueness of LLM training process and gives top priorities to detecting and recovering failures in a routine manner. Leveraging parallelisms and characteristics of LLM training, ByteRobust enables high-capacity fault tolerance, prompt fault demarcation, and localization with an effective data-driven approach, comprehensively ensuring continuous and efficient training of LLM tasks. ByteRobust is deployed on a production GPU platform with over 200,000 GPUs and advances the state of the art in training robustness by achieving 97% ETTR for a three-month training job on 9,600 GPUs. Borui Wan, Gaohong Liu, Zuquan Song, Jun Wang 0039, Guangming Sheng, Shuguang Wang, Houmin Wei, Weiqiang Lou, Mofan Zhang, Kaihua Jiang, Cheng Ren, Xiaoyun Zhi, Menghan Yu, Zhe Nan, Zhuolin Zheng, Baoquan Zhong, Qinlong Wang, Jinxin Chi, Wang Zhang 0017, Zixian Du, Sida Zhao, Jingzhe Tang, Zherui Liu, Chuan Wu 0001, Yanghua Peng, Haibin Lin, Wencong Xiao, Xin Liu 0086 |
SOSP | 7 |
| 2025 | Radio Frequency Fingerprinting for WiFi Authentication Based on Detrended Fluctuation AnalysisabstractTraditional physical‐layer authentication (PLA) approaches primarily rely on a limited set of hardware features, which restricts their robustness and accuracy in dynamic wireless environments. This paper proposes a novel PLA enhancement framework based on nonlinear signal analysis, introducing detrended fluctuation analysis (DFA) as a distinctive hardware fingerprinting feature. We first model the statistical properties of DFA and analytically derive its dependence on intrinsic hardware imperfections, thereby establishing its feasibility for device authentication. To improve overall system performance, the DFA feature is further fused with conventional features such as fractal dimension and carrier frequency offset (CFO) within a shallow classification framework. Experimental evaluation is conducted on real signal data collected from 28 commercial devices, demonstrating that the proposed multifeature PLA scheme can significantly improve authentication accuracy, confirming the effectiveness of DFA in enhancing physical‐layer security without additional cryptographic overhead. Shuguang Wang, Songyan Li, Tingjia Liu, Shuangrui Zhao, Yulong Shen 0001 |
IET Inf. Secur. | 1 |
| 2024 | Physical layer authentication in the internet of vehicles through multiple vehicle-based physical attributes prediction
Mubarak Umar, Shuguang Wang, Minggang Zheng, Zhiwei Zhang 0004, Yulong Shen 0001 |
Ad Hoc Networks | 4 |
| 2021 | WILSON: A Divide and Conquer Approach for Fast and Effective News Timeline Summarization
Yiming Liao, Shuguang Wang, Dongwon Lee 0001 |
EDBT | 2 |
| 2019 | Characterization and Early Detection of Evergreen News Articles
Yiming Liao, Shuguang Wang, Eui-Hong Han, Jongwuk Lee, Dongwon Lee 0001 |
ECML/PKDD (3) | 2 |
| 2016 | Predicting the shape and peak time of news article viewsabstractPredicting the popularity of news articles - whether measured via retweets, clicks, or views - is an important problem for editors, journalists, and readers alike. In this paper, we introduce a new model to predict the shape of news article views, and use this model to determine when an article will likely reach its maximum number of views. Although volume prediction for news articles has been extensively studied predicting when a burst of views will happen, in what shape, and by how much, remains an open problem. We engineer several classes of features (metadata, contextual or content-based, temporal, and social), develop models to classify shape of views, with particular attention paid to performing online, time-updated, prediction, i.e., using data before and during the early stages of article prediction to predict its eventual peak views and update earlier predictions. The system presented here is an emerging application being developed at The Washington Post and can be used to support article placement, updating, and promotion strategies. Yaser Keneshloo, Shuguang Wang, Eui-Hong Han, Naren Ramakrishnan |
IEEE BigData | 2 |
| 2016 | Predicting the Popularity of News ArticlesabstractConsuming news articles is an integral part of our daily lives and news agencies such as The Washington Post (WP) expend tremendous effort in providing high quality reading experiences for their readers. Journalists and editors are faced with the task of determining which articles will become popular so that they can efficiently allocate resources to support a better reading experience. The reasons behind the popularity of news articles are typically varied, and might involve contemporariness, writing quality, and other latent factors. In this paper, we cast the problem of popularity prediction problem as regression, engineer several classes of features (metadata, contextual or content-based, temporal, and social), and build models for forecasting popularity. The system presented here is deployed in a real setting at The Washington Post; we demonstrate that it is able to accurately predict article popularity with an R2 ≈ 0.8 using features harvested within 30 minutes of publication time. Yaser Keneshloo, Shuguang Wang, Eui-Hong Han, Naren Ramakrishnan |
SDM | 2 |
| 2015 | BreakFast: Analyzing Celerity of NewsabstractIn the hypercompetitive news market, news outlets race to break news first. In order to provide better breaking news service and improve the reader experience, news agencies need to understand how to identify bottlenecks and streamline their reporting and delivery processes. With that in mind, we built a system, BreakFast, to measure and compare the speed of delivery of breaking news from various news sources to readers. One of the primary challenges of this comparison is how to identify which breaking news items are about the same emerging event but reported by different news agencies with different headlines and content. To tackle this problem, we extracted keywords automatically from the content, identified important topics, and then developed a classification model. The model identifies the same breaking stories from multiple news sources with an accuracy of approximately 90%. We also proposed new metrics to evaluate the speed of breaking news services and built real-time dashboards to monitor performance over time. We deployed BreakFast into the breaking news service at The Washington Post. This integrated system narrowed in on bottlenecks in its breaking news generation and delivery process, and improved its breaking news service in terms of time by more than 50%. Shuguang Wang, Eui-Hong Han |
ICMLA | 1 |
| 2013 | Global and local information combined to detect singular points in fingerprint images
Shuguang Wang, Tiande Guo |
Sci. China Inf. Sci. | 2 |
| 2012 | A Novel Task Management System for Modelica-Based Multi-discipline Virtual Experiment PlatformabstractCurrently, there is few uniform modelling standards for virtual experiment (VE) systems of different disciplines. The scalability and compatibility of existing systems are relatively poor. The idea of Modelica provides a good opportunity for the unification of the modelling of multi-discipline VEs (MDVE). However, Modelica is an original multi-domain modelling method for scientific research, instead of for VE education. There are some obvious gaps to bring it into a MDVE platform (MDVEP), especially, if it is designed to support massive users and parallel modelling and resolving. This paper presents a new virtual experiment distributed task management system (VETMS) for MDVEP, which can improve the efficiency, stability and availability of the platform. It uses hierarchical design method to decouple the different modules and also provides a set of Application Programming Interfaces (APIs) for external calls. Besides, the system performance is also considered. A Modelica-oriented mechanism is proposed to tackle service failures. A parallelized Twisted framework is presented to overcome the problem of limitation of concurrent requests. Meanwhile, a NAT (Network Address Translator) traversal module of TCP based STUNT protocol is added to reduce the amount of data through master node. Experiment results show that the system can serve as a task management service for MDVEP with good performance. Wenbin Jiang 0001, Shuguang Wang, Hai Jin 0001 |
APSCC | 2 |
| 2012 | Keyword annotation of biomedicai documents with graph-based similarity methodsabstractIn this paper, we present a new approach that lets us extract, and represent relations among terms (concepts) in the documents and uses these relations to support various document analysis applications. Our approach works by building a graph of local co-occurrence relations among terms that are extracted directly from text and by defining a global similarity metric among these terms and sets of terms using the graph and its connectivity. We demonstrate the benefit of the approach on the problem of MeSH keyword annotation of documents based on their abstracts. Shuguang Wang, Milos Hauskrecht |
BIBM | 1 |
| 2011 | An Efficient Framework for Constructing Generalized Locally-Induced Text Metrics
Saeed Amizadeh, Shuguang Wang, Milos Hauskrecht |
IJCAI | 2 |
| 2010 | Effective query expansion with the resistance distance based term similarity metricabstractIn this paper, we define a new query expansion method that relies on term similarity metric derived from the electric resistance network. This proposed metric lets us measure the mutual relevancy in between terms and between their groups. This paper shows how to define this metric automatically from the document collection, and then apply it in query expansion for document retrieval tasks. The experiments show this method can be used to find good expansion terms of search queries and improve document retrieval performance on two TREC genomic track datasets. Shuguang Wang, Milos Hauskrecht |
SIGIR | 1 |
| 2008 | Improving biomedical document retrieval using domain knowledgeabstractResearch articles typically introduce new results or findings and relate them to knowledge entities of immediate relevance. However, a large body of context knowledge related to the results is often not explicitly mentioned in the article. To overcome this limitation the state-of-the-art information retrieval approaches rely on the latent semantic analysis in which terms in articles are projected to a lower dimensional latent space and best possible matches in this space are identified. However, this approach may not perform well enough if the number of explicit knowledge entities in the articles is too small compared to the amount of knowledge in the domain. We address the problem by exploiting a domain knowledge layer, a rich network of relations among knowledge entities in the domain extracted from a large corpus of documents. The knowledge layer supplies the context knowledge that lets us relate different knowledge entities and hence improve the information retrieval performance. We develop and study a new framework for i) learning and aggregating the relations in the knowledge layer from the literature corpus; ii) and for exploiting these relations to improve the information-retrieval of relevant documents. Shuguang Wang, Milos Hauskrecht |
SIGIR | 1 |
| 2008 | Singular Points Detection Based on Zero-Pole Model in Fingerprint ImagesabstractAn algorithm is proposed which combines Zero-pole Model and Hough Transform(HT) to detect singular points. Orientation of singular points is defined on basis of Zero-pole Model which can further explain the practicability of Zero-pole Model. Contrary to orientation field generation, detection of singular points is simplified to determine the parameters of Zero-pole Model. HT uses rather global information of fingerprint images to detect singular points. This makes our algorithm more robust to noise than methods which only use local information. As Zero-pole Model may have a little warp from actual fingerprint orientation field, Poincare index is used to make position adjustment in neighborhood of the detected candidate singular points. Experimental results show that our algorithm performs well and fast enough for real time application in database NIST-4. Shuguang Wang, Hongfa Wang, Tiande Guo |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2007 | Efficient index-based KNN join processing for high-dimensional data
Cui Yu, Bin Cui 0001, Shuguang Wang, Jianwen Su |
Inf. Softw. Technol. | 3 |
| 2003 | Structured use of external knowledge for event-based open domain question answeringabstractOne of the major problems in question answering (QA) is that the queries are either too brief or often do not contain most relevant terms in the target corpus. In order to overcome this problem, our earlier work integrates external knowledge extracted from the Web and WordNet to perform Event-based QA on the TREC-11 task. This paper extends our approach to perform event-based QA by uncovering the structure within the external knowledge. The knowledge structure loosely models different facets of QA events, and is used in conjunction with successive constraint relaxation algorithm to achieve effective QA. Our results obtained on TREC-11 QA corpus indicate that the new approach is more effective and able to attain a confidence-weighted score of above 80%. Grace Hui Yang, Tat-Seng Chua, Shuguang Wang, Chun-Keat Koh |
SIGIR | 3 |
| 2001 | Compressing the Index - A Simple and yet Efficient Approximation Approach to High-Dimensional Indexing
Shuguang Wang, Cui Yu, Beng Chin Ooi |
WAIM | 1 |