EDBT 2026 Demo / reviewers in the wild / expert
Shiming Yang
dblp:30/3942
· DBLP profile ↗
11ranked-venue papers
7as first author
6since 2021 · last 2026
0000-0003-0338-4268ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 4 · 3 first-author · 1 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1Computer networks · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Language models and text generation · 67% Reinforcement learning · 33% | |
| Databases, data mining, and information retrieval
1 paper |
Graph data management · 67% Database system architecture and tuning · 33% |
Topics — the 5 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
chain-of-thought reasoning |
0.9 | 1 | 2025 | Demystifying Long Chain-of-Thought Reasoning · ICML 2025 |
Natural language and speech › Language models and text generation › chain-of-thought reasoning
long chain-of-thought reasoning |
0.9 | 1 | 2025 | Demystifying Long Chain-of-Thought Reasoning · ICML 2025 |
Machine learning › Reinforcement learning › reinforcement learning for NLP
reinforcement learning for language models |
0.9 | 1 | 2025 | Demystifying Long Chain-of-Thought Reasoning · ICML 2025 |
Graph data management
graph database |
0.9 | 1 | 2025 | An efficient and scalable graph database with built-in temporal support · VLDB J. 2025 |
Graph data management › graph database
temporal graph database |
0.9 | 1 | 2025 | An efficient and scalable graph database with built-in temporal support · VLDB J. 2025 |
Methods — techniques the papers use, named apart from their topics
temporal indexing · 0.9supervised fine-tuning · 0.9reward shaping · 0.9reinforcement learning · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Doux: Decoupling Values from Keys for Real-Time Analytics
Shiming Yang, Shuang Liu 0007, Wei Lu 0015, Kuien Liu, Yuxing Chen 0003, Anqun Pan, Lixiong Zheng, Xiaoyong Du 0001 |
ICDE | 1 |
| 2025 | Demystifying Long Chain-of-Thought ReasoningabstractScaling inference compute has become a key driver of advanced reasoning in large language models (LLMs). A proven approach for scaling inference compute is to generate long chains-of-thought (CoTs), enabling models to engage in structured reasoning strategies such as backtracking and error correction. Reinforcement learning (RL) has emerged as a crucial method for developing these capabilities, yet the conditions under which long CoTs emerge remain unclear, and RL training requires careful design choices. In this study, we systematically investigate the underlying mechanics of long CoT reasoning—examining the factors that enable models to generate extended reasoning trajectories. Through extensive supervised fine-tuning (SFT) and RL experiments, we identify three key findings: 1) while SFT is not strictly necessary, it significantly simplifies training and improves efficiency; 2) reasoning capabilities tend to emerge with increased training compute but are not guaranteed, making reward shaping essential for stabilizing CoT length growth; and 3) scaling verifiable reward signals is critical for RL, and we find that leveraging noisy, web-extracted solutions with filtering mechanisms shows promising potential, particularly in out-of-distribution (OOD) reasoning tasks such as STEM problem-solving. These insights provide practical guidance for optimizing training strategies to enhance long CoT reasoning in LLMs. Shiming Yang, Yuxuan Tong, Xinyao Niu, Graham Neubig, Xiang Yue |
ICML | 1 |
| 2025 | An efficient and scalable graph database with built-in temporal support
Jiamin Hou, Zhanhao Zhao, Wei Lu 0015, Shiming Yang, Shuang Liu 0007, Quanqing Xu, Chuanhui Yang, Xiaoyong Du 0001 |
VLDB J. | 4 |
| 2023 | Active Reconfigurable Intelligent Surface-Aided Cognitive Radio SystemabstractThis paper considers an active reconfigurable intelligent surface (RIS)-aided multiple-input single-output (MISO) cognitive radio (CR) system where the active RIS is designed to assist the transmission of the secondary user (SU) by tuning its controllable reflecting elements (REs) with enhanced reflections. As each RE in active RISs can adjust the phases and amplify the incident signals, active RISs thus provide a more favorable channel condition for the considered CR system than conventional passive RISs do. We aim to maximize the achievable rate of the SU subject to the interference temperature (IT) constraints on the primary users (PUs) as well as the power budget constraint on the active RIS. Towards this end, we propose an alternating optimization algorithm to solve this rate maximization problem. More specifically, the transmit beamforming vector is obtained by solving a second-order core programming (SOCP) problem, and the reflect beamforming vector is obtained with the aid of the fractional programming (FP) technique. Numerical results are provided to compare the active RIS-aided CR system with the passive RIS-aided one, showing that under the same power budget, the active RIS notably outperforms the passive RIS when they are used to assist the transmission of the SU in the CR system. Shiming Yang, Ruizhe Long, Ying-Chang Liang |
ICC | 1 |
| 2022 | Incomplete multi-view clustering with multiple imputation and ensemble clustering
Guoqing Chao, Shiming Yang, Chunshan Li |
Appl. Intell. | 3 |
| 2021 | The prognostic value of an autophagy-related lncRNA signature in hepatocellular carcinomaabstractBACKGROUND: lncRNA may be involved in the occurrence, metastasis, and chemical reaction of hepatocellular carcinoma (HCC) through various pathways associated with autophagy. Therefore, it is urgent to reveal more autophagy-related lncRNAs, explore these lncRNAs' clinical significance, and find new targeted treatment strategies. METHODS: The corresponding data of HCC patients and autophagy genes were obtained from the TCGA database, and the human autophagy database respectively. Based on the co-expression and Cox regression analysis to construct prognostic prediction signature. RESULTS: Finally, a signature containing seven autophagy-related lncRNAs (PRRT3-AS1, RP11-479G22.8, RP11-73M18.8, LINC01138, CTD-2510F5.4, CTC-297N7.9, RP11-324I22.4) was constructed. Based on the risk score of signature, Overall survival (OS) curves show that the OS of high-risk patients is significantly lower than that of low-risk patients (P = 2.292e-10), and the prognostic prediction accuracy of risk score (AUC = 0.786) is significantly higher than that of ALBI (0.532), child_pugh (0.573), AFP (0.5751), and AJCC_stage (0.631). Moreover, multivariate Cox analysis and Nomogram of risk score are indicated that the 1-year and 3-year survival rates of patients are obviously accuracy by the combined analysis of the risk score, child_pugh, age, M_stage, and Grade (The AUC of 1- and 3-years are 0.87, and 0.855). Remarkably, the 7 autophagy-related lncRNAs may participate in Spliceosome, Cell cycle, RNA transport, DNA replication, and mRNA surveillance pathway and be related to the biological process of RNA splicing and mRNA splicing. CONCLUSION: In conclusion, the 7 autophagy-related lncRNAs might be promising prognostic and therapeutic targets for HCC. Shiming Yang, Yaping Zhou, Xiangxin Zhang, Jianfeng Fu |
BMC Bioinform. | 1 |
| 2014 | Design of Vendor-neutral Platform for Fast Prototype Model Verification and Deployment
Shiming Yang, Peter Fu-Ming Hu, Yulei Wang 0002, Amechi N. Anazodo, Catriona Miller, Raymond Fang, Stacy Shackelford, Colin F. Mackenzie |
AMIA | 1 |
| 2014 | Detecting Road Traffic Events by Coupling Multiple Timeseries With a Nonparametric Bayesian MethodabstractRoad traffic sensors provide rich multivariable datastreams about the current traffic conditions. Occasionally, there are unusual traffic events (such as accidents, jams, and severe weather) that disrupt the expected road traffic conditions. Detecting the occurrence of such events in an online and real-time manner is useful to drivers in planning their routes and in the management of the transportation infrastructure. We propose a new method for detecting traffic events that impact road traffic conditions by extending the Bayesian robust principal component analysis (RPCA) approach. Our method couples multiple traffic datastreams so that they share a certain sparse structure. This sparse structure is used to localize traffic events in space and time. The traffic datastreams are measurements of different physical quantities (e.g., traffic flow and road occupancy) by different nearby sensors. Our proposed method processes datastreams in an incremental way with small computational cost; hence, it is suitable to detect events in an online and real-time manner. We experimentally analyze the detection performance of the proposed coupled Bayesian RPCA (BRPCA) using real data from loop detectors on the Minnesota I-494. We find that our method significantly improves the detection accuracy when compared with the traditional PCA and noncoupled BRPCA. Shiming Yang, Konstantinos Kalpakis, Alain Biem |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2013 | Real-world Respiratory Rate (RR) Signal Processing During Trauma Patient Resuscitation
R. North, Peter Fu-Ming Hu, Shiming Yang, K. Frank, Colin F. Mackenzie |
AMIA | 3 |
| 2012 | Online Recovery of Missing Values in Vital Signs Data Streams Using Low-Rank Matrix CompletionabstractContinuous, automated, electronic patient vital signs data are important to physicians in evaluating traumatic brain injury (TBI) patients' physiological status and reaching timely decisions for therapeutic interventions. However, missing values in the medical data streams hinder applying many standard statistical or machine learning algorithms and result in losing some episodes of clinical importance. In this paper, we present a novel approach to filling missing values in streams of vital signs data. We construct sequences of Hankel matrices from vital signs data streams, find that these matrices exhibit low-rank, and utilize low-rank matrix completion methods from compressible sensing to fill in the missing data. We demonstrate that our approach always substantially outperforms other popular fill-in methods, like k-nearest-neighbors and expectation maximization. Further, we show that our approach recovers thousands of simulated missing data for intracranial pressure, a critical stream of measurements for guiding clinical interventions and monitoring traumatic brain injuries. Shiming Yang, Konstantinos Kalpakis, Colin F. Mackenzie, Lynn G. Stansbury, Deborah M. Stein, Thomas M. Scalea, Peter Fu-Ming Hu |
ICMLA (1) | 1 |
| 2012 | Performance study on CUDA GPUs for parallelizing the local ensemble transformed Kalman filter algorithmabstractSUMMARY Modern graphics cards provide computational capabilities that exceed current CPUs. As one of the computational intensive problems, numerical weather prediction has the opportunity to benefit from the massive number of threads and large memory throughput in the graphics architecture. In this paper, we present the key steps to integrate the Compute Unified Device Architecture (CUDA) programming framework for one key component in numerical weather prediction, the data assimilation algorithm, which incorporates the observational data into the model to produce the best initial condition in the next prediction. The data assimilation algorithm we studied in this paper exhibits good localization and favors parallelism. To maximize the throughput of the graphics card, over a million CUDA threads, global memory coalescing, and fast graphics shared memory are utilized. We also demonstrate the differences in the advancement of GPU architectures from the GTX 200 series to Fermi. The experiments are carried out separately on a GTX 260 (GTX 200 series) and a GTX 460 (Fermi) graphics card. Results show an improvement of 72.1× speedup running on the GTX 260 and 92.7×speedup on the GTX 460. The results provide attractive evidence for applying CUDA GPUs to high demanding scientific computation realms. Copyright © 2011 John Wiley & Sons, Ltd. Timothy Blattner, Shiming Yang |
Concurr. Comput. Pract. Exp. | 2 |