Senmiao Wang

dblp:237/8191 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
5since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Theoretical computer science
2 papers
Mathematical optimization · 100%
Artificial intelligence
2 papers
Language models and text generation · 52% Optimization for machine learning · 26% Reinforcement learning · 23%

Topics — the 11 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Optimization for machine learning
constrained optimization
1.012026
VCORE: Variance-Controlled Optimization-based Reweighting for Chain-of-Thought Supervision · ACL (1) 2026
Natural language and speech › Language models and text generation
large language model fine-tuning
1.012026
VCORE: Variance-Controlled Optimization-based Reweighting for Chain-of-Thought Supervision · ACL (1) 2026
Natural language and speech › Language models and text generation
token reweighting
1.012026
VCORE: Variance-Controlled Optimization-based Reweighting for Chain-of-Thought Supervision · ACL (1) 2026
Mathematical optimization › continuous optimization › convex optimization
first-order methods
0.812024
PDHG-Unrolled Learning-to-Optimize Method for Large-Scale Linear Programming · ICML 2024
Mathematical optimization › linear programming
large-scale linear programming
0.812024
PDHG-Unrolled Learning-to-Optimize Method for Large-Scale Linear Programming · ICML 2024
Mathematical optimization
linear programming
0.812024
PDHG-Unrolled Learning-to-Optimize Method for Large-Scale Linear Programming · ICML 2024
Mathematical optimization › primal-dual method
primal-dual hybrid gradient
0.812024
PDHG-Unrolled Learning-to-Optimize Method for Large-Scale Linear Programming · ICML 2024
Mathematical optimization › optimization for machine learning
unrolled optimization
0.812024
PDHG-Unrolled Learning-to-Optimize Method for Large-Scale Linear Programming · ICML 2024
Machine learning › Reinforcement learning
markov decision process
0.612022
Adaptive Model Design for Markov Decision Process · ICML 2022
Mathematical optimization
bilevel optimization
0.612022
Adaptive Model Design for Markov Decision Process · ICML 2022
Machine learning › Reinforcement learning
reinforcement learning from human feedback
0.312026
VCORE: Variance-Controlled Optimization-based Reweighting for Chain-of-Thought Supervision · ACL (1) 2026

Methods — techniques the papers use, named apart from their topics

iterative prediction · 1.1bilevel programming · 1.1variance control · 1.0supervised fine-tuning · 1.0constrained optimization · 1.0neural network unrolling · 0.8graph neural network channel expansion · 0.8PDHG · 0.8
YearPublicationVenuePosition
2026 VCORE: Variance-Controlled Optimization-based Reweighting for Chain-of-Thought Supervision
abstract
Supervised fine-tuning (SFT) on long chainof-thought (CoT) trajectories has emerged as a crucial technique for enhancing the reasoning abilities of large language models (LLMs).However, the standard cross-entropy loss treats all tokens equally, ignoring their heterogeneous contributions across a reasoning trajectory.This uniform treatment leads to misallocated supervision and weak generalization, especially in complex, long-form reasoning tasks.To address this, we introduce Variance-Controlled Optimization-based REweighting (VCORE), a principled framework that reformulates CoT supervision as a constrained optimization problem.By adopting an optimization-theoretic perspective, VCORE enables a principled and adaptive allocation of supervision across tokens, thereby aligning the training objective more closely with the goal of robust reasoning generalization.Empirical evaluations demonstrate that VCORE achieves the strongest overall average performance, with especially clear gains on lower-capacity models.Across both in-domain and out-of-domain settings, VCORE achieves substantial performance gains on mathematical and coding benchmarks, using models from the Qwen3 series (4B, 8B, 32B) and LLaMA-3.1-8B-Instruct.Moreover, we show that VCORE serves as a more effective initialization for subsequent reinforcement learning, establishing a stronger foundation for advancing the reasoning capabilities of LLMs. 1
Senmiao Wang, Hanbo Huang, Ruoyu Sun 0001, Shiyu Liang
ACL (1)2
2024 PDHG-Unrolled Learning-to-Optimize Method for Large-Scale Linear Programming
abstract
Solving large-scale linear programming (LP) problems is an important task in various areas such as communication networks, power systems, finance and logistics. Recently, two distinct approaches have emerged to expedite LP solving: (i) First-order methods (FOMs); (ii) Learning to optimize (L2O). In this work, we propose an FOM-unrolled neural network (NN) called PDHG-Net, and propose a two-stage L2O method to solve large-scale LP problems. The new architecture PDHG-Net is designed by unrolling the recently emerged PDHG method into a neural network, combined with channel-expansion techniques borrowed from graph neural networks. We prove that the proposed PDHG-Net can recover PDHG algorithm, thus can approximate optimal solutions of LP instances with a polynomial number of neurons. We propose a two-stage inference approach: first use PDHG-Net to generate an approximate solution, and then apply PDHG algorithm to further improve the solution. Experiments show that our approach can significantly accelerate LP solving, achieving up to a 3$\times$ speedup compared to FOMs for large-scale LP problems.
Bingheng Li, Linxin Yang, Senmiao Wang, Haitao Mao, Yao Ma 0001, Akang Wang, Tian Ding, Jiliang Tang, Ruoyu Sun 0001
ICML4
2022 Adaptive Model Design for Markov Decision Process
abstract
In a Markov decision process (MDP), an agent interacts with the environment via perceptions and actions. During this process, the agent aims to maximize its own gain. Hence, appropriate regulations are often required, if we hope to take the external costs/benefits of its actions into consideration. In this paper, we study how to regulate such an agent by redesigning model parameters that can affect the rewards and/or the transition kernels. We formulate this problem as a bilevel program, in which the lower-level MDP is regulated by the upper-level model designer. To solve the resulting problem, we develop a scheme that allows the designer to iteratively predict the agent’s reaction by solving the MDP and then adaptively update model parameters based on the predicted reaction. The algorithm is first theoretically analyzed and then empirically tested on several MDP models arising in economics and robotics.
Siyu Chen 0001, Donglin Yang, Jiayang Li 0001, Senmiao Wang, Zhuoran Yang, Zhaoran Wang 0001
ICML4
2022 KRTunnel: DNS channel detector for mobile devices
abstract
Nowadays, DNS channel attacks on mobile devices have become a challenging threat. Attackers usually attack mobile devices and steal information with the help of DNS channel. It is difficult for users to detect this kind of attack, especially when attackers covert sensitive information in the DNS response. In this paper, we proposed a method for DNS tunnel detection based on isolated forest for Android. We constructed a framework for mobile devices to collect DNS tunnel traffic. Based on the analysis of DNS tunnel traffic generated on mobile devices, we extracted features based on DNS request and response and constructed the feature set. We proposed a DNS tunnel detector, KRTunnel, for mobile devices. Experiments showed that KRTunnel can identify unseen DNS tunnel traffic with the accuracy of 98.1%.
Senmiao Wang, Luli Sun, Su-Juan Qin, Wenmin Li 0001
Comput. Secur.1
2022 KRProtector: Detection and Files Protection for IoT Devices on Android Without ROOT Against Ransomware Based on Decoys
abstract
Nowadays, cryptographic ransomware on Android has become one of the most serious threat. They extort users by means of encrypting private data on their devices. Even worse, there exists little files protection solution on IoT devices without ROOT. In light of this, there is an urgent need for countermeasure solutions on IoT devices without ROOT. In this article, we analyze characteristics of cryptographic ransomware. We propose the strategy of files protection against ransomware based on decoys. In order to satisfy the need of files protection on devices without root, we design and implement KRProtector to detect ransomware and protect files based on decoys.
Senmiao Wang, Hua Zhang 0001, Su-Juan Qin, Wenmin Li 0001, Tengfei Tu, Ana Shen
IEEE Internet Things J.1
2020 A Multiclass Detection System for Android Malicious Apps Based on Color Image Features
abstract
The visual recognition of Android malicious applications (Apps) is mainly focused on the binary classification using grayscale images, while the multiclassification of malicious App families is rarely studied. If we can visualize the Android malicious Apps as color images, we will get more features than using grayscale images. In this paper, a method of color visualization for Android Apps is proposed and implemented. Based on this, combined with deep learning models, a multiclassifier for the Android malicious App families is implemented, which can classify 10 common malicious App families. In order to better understand the behavioral characteristics of malicious Apps, we conduct a comprehensive manual analysis for a large number of malicious Apps and summarize 1695 malicious behavior characteristics as customized features. Compared with the App classifier based on the grayscale visualization method, it is verified that the classifier using the color visualization method can achieve better classification results. We use four types of Android App features: classes.dex file, sets of class names, APIs, and customized features as input for App visualization. According to the experimental results, we find out that using the customized features as the color visualization input features can achieve the highest detection accuracy rate, which is 96% in the ten malicious families.
Hua Zhang 0001, Jiawei Qin, Boan Zhang, Fei Gao 0001, Senmiao Wang, Yangye Hu
Wirel. Commun. Mob. Comput.7