EDBT 2026 Demo / reviewers in the wild / expert
Yanhang Zhang
dblp:227/6681
· DBLP profile ↗
7ranked-venue papers
3as first author
7since 2021 · last 2025
0009-0002-1531-665XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Theory of computation · 3 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Rethinking Hard Thresholding Pursuit: Full Adaptation and Sharp EstimationabstractHard Thresholding Pursuit (HTP) has aroused increasing attention for its robust theoretical guarantees and impressive numerical performance in non-convex optimization. This paper consider a high-dimensional linear regression model withnobservations,ppredictors, and an unknowns∗-sparse signal β∗∈ Rpcorrupted by noise of magnitude σ.We introduce a novel tuning-free procedure, namely Full-Adaptive HTP (FAHTP), that simultaneously adapts to both the unknown sparsity and signal strength of the underlying model. Our theoretical analysis rigorously characterizes the iterative thresholding dynamics of FAHTP, offering refined theoretical insights. In specific, under the beta-min condition min{i:β∗i̸=0}|β∗i| ≥ Cσ(logp/n)1/2, FAHTP achieves oracle estimation rate σ(s∗/n)1/2, highlighting its theoretical superiority over convex competitors such as LASSO and SLOPE, and recovers the true support set exactly. More importantly, even without the beta-min condition, FAHTP achieves a tighter error bound than the classical minimax rate with high probability. The comprehensive numerical experiments substantiate our theoretical findings, underscoring the effectiveness and robustness of the proposed FAHTP. Yanhang Zhang, Shixiang Liu, Zhifan Li, Jianxin Yin 0001 |
IEEE Trans. Inf. Theory | 1 |
| 2024 | A minimax optimal approach to high-dimensional double sparse linear regressionabstractIn this paper, we focus our attention on the high-dimensional double sparse linear regression, that is, a combination of element-wise and group-wise sparsity. To address this problem, we propose an IHT-style (iterative hard thresholding) procedure that dynamically updates the threshold at each step. We establish the matching upper and lower bounds for parameter estimation, showing the optimality of our proposal in the minimax sense. More importantly, we introduce a fully adaptive optimal procedure designed to address unknown sparsity and noise levels. Our adaptive procedure demonstrates optimal statistical accuracy with fast convergence. Additionally, we elucidate the significance of the element-wise sparsity level $s_0$ as the trade-off between IHT and group IHT, underscoring the superior performance of our method over both. Leveraging the beta-min condition, we establish that our IHT-style procedure can attain the oracle estimation rate and achieve almost full recovery of the true support set at both the element level and group level. Finally, we demonstrate the superiority of our method by comparing it with several state-of-the-art algorithms on both synthetic and real-world datasets. Yanhang Zhang, Zhifan Li, Shixiang Liu, Jianxin Yin 0001 |
J. Mach. Learn. Res. | 1 |
| 2024 | Estimating Double Sparse Structures Over ℓu(ℓq) -Balls: Minimax Rates and Phase TransitionabstractIn this paper, we focus on the high-dimensional double sparse structures, where the parameter of interest simultaneously encourages group-wise and element-wise sparsity. By combining the Gilbert-Varshamov bound and its variants, we develop a novel lower bound technique for the metric entropy of the parameter space, specifically tailored for the double sparse structure over$\ell _{u}(\ell _{q})$-balls with$u,q \in [0,2$). We give lower bounds on the estimation error using an information-theoretic approach, leveraging the proposed technique and Fano’s inequality. To complement the lower bounds, we establish matching upper bounds through a direct analysis of constrained least-squares estimators and utilizing results from empirical processes. A significant discovery is that a phase transition phenomenon exists on the minimax rates for$u,q \in (0, 2)$. Furthermore, we extend the theoretical findings to the double sparse regression models and determine the minimax rates for estimation error. A novel Double Sparse Iterative Hard Thresholding (DSIHT) procedure is developed, with minimax optimality guaranteed. Finally, we demonstrate the superiority of the proposed method through numerical experiments. Zhifan Li, Yanhang Zhang, Jianxin Yin 0001 |
IEEE Trans. Inf. Theory | 2 |
| 2023 | A Splicing Approach to Best Subset of Groups SelectionabstractBest subset of groups selection (BSGS) is the process of selecting a small part of nonoverlapping groups to achieve the best interpretability on the response variable. It has attracted increasing attention and has far-reaching applications in practice. However, due to the computational intractability of BSGS in high-dimensional settings, developing efficient algorithms for solving BSGS remains a research hotspot. In this paper, we propose a group-splicing algorithm that iteratively detects the relevant groups and excludes the irrelevant ones. Moreover, coupled with a novel group information criterion, we develop an adaptive algorithm to determine the optimal model size. Under certain conditions, it is certifiable that our algorithm can identify the optimal subset of groups in polynomial time with high probability. Finally, we demonstrate the efficiency and accuracy of our methods by comparing them with several state-of-the-art algorithms on both synthetic and real-world data sets. History: Accepted by Andrea Lodi, Area Editor for Design & Analysis of Algorithms-Discrete. Funding: This work was supported by National Natural Science Foundation of China [Grants 72171216, 71921001, and 71991474], the Key Research and Development Program of Guangdong [Grant 2019B020228001], the Science and Technology Program of Guangzhou, China [Grant 202002030129], The Fundamental Research Funds for the Central Universities and the Research Funds of Renmin University of China [Grant 22XNH161], and the Outstanding Graduate Student Innovation and Development Program of Sun Yat-Sen University [Grant 19lgyjs64]. Supplemental Material: The online appendix and video are available at https://doi.org/10.1287/ijoc.2022.1241 . Yanhang Zhang, Junxian Zhu |
INFORMS J. Comput. | 1 |
| 2022 | V-Doc : Visual questions answers with DocumentsabstractWe propose V-Doc, a question-answering tool using document images and PDF, mainly for researchers and general non-deep learning experts looking to generate, process, and understand the document visual question answering tasks. The V-Doc supports generating and using both extractive and abstractive question-answer pairs using documents images. The extractive QA selects a subset of tokens or phrases from the document contents to predict the answers, while the abstractive QA recognises the language in the content and generates the answer based on the trained model. Both aspects are crucial to understanding the documents, especially in an image format. We include a detailed scenario of question generation for the abstractive QA task. V-Doc supports a wide range of datasets and models, and is highly extensible through a declarative, framework-agnostic platform.11Data and demo video: https://github.com/usydnlp/vdoc Yihao Ding, Runlin Wang, Yanhang Zhang, Xianru Chen, Yuzhong Ma, Hyunsuk Chung, Soyeon Caren Han |
CVPR | 4 |
| 2022 | Equivalent Mutants Detection Based on Weighted Software Behavior GraphabstractThe equivalent mutants problem is one of the crucial problems in mutation testing. In consequence of its existence, the effectiveness of mutation testing is underestimated. In addition, it will produce a certain amount of useless overhead. Equivalent mutants cannot be detected by any test input. The existing works mostly focus on static analysis to detect, or avoid generating, the equivalent mutants. The essence of these methods is to use prior knowledge to establish some rules of program equivalence. However, (1) it needs a lot of professional labor to sort out the equivalence rules, and (2) only a small part of the rules can be determined in advance, because of the diversity of mutation operators and mutation targets. Consequently, the best result reported so far is 50% of the equivalent mutants can be detected. Since it is generally believed that manual judgment of program equivalence is the most reliable, this paper proposes a novel method to automatically detect equivalent mutants by tracing program behavior like the professionals. The weighted software behavior graph is utilized in the detection of equivalent mutants for the first time. This method can not only figure out different execution paths, but also be sensitive to execution frequency. By comparing the weighted software behavior graphs of an alive mutant and its original program, we are able to examine more precisely whether the alive mutant is the same as the original program, in terms of the state of infection and/or the propagation. Evaluation results on an open dataset of manually evaluated equivalent mutants show that our approach can detect 77.5% of all the equivalent mutants, which is much higher than the existing static methods. Dan Gong, Tiantian Wang 0001, Xiaohong Su, Yanhang Zhang |
Int. J. Softw. Eng. Knowl. Eng. | 4 |
| 2022 | abess: A Fast Best-Subset Selection Library in Python and RabstractWe introduce a new library named abess that implements a unified framework of best-subset selection for solving diverse machine learning problems, e.g., linear regression, classification, and principal component analysis. Particularly, abess certifiably gets the optimal solution within polynomial time with high probability under the linear model. Our efficient implementation allows abess to attain the solution of best-subset selection problems as fast as or even 20x faster than existing competing variable (model) selection toolboxes. Furthermore, it supports common variants like best subset of groups selection and $\ell_2$ regularized best-subset selection. The core of the library is programmed in C++. For ease of use, a Python library is designed for convenient integration with scikit-learn, and it can be installed from the Python Package Index (PyPI). In addition, a user-friendly R library is available at the Comprehensive R Archive Network (CRAN). The source code is available at: https://github.com/abess-team/abess. Liyuan Hu, Kangkang Jiang, Yanhang Zhang, Shiyun Lin, Junxian Zhu |
J. Mach. Learn. Res. | 6 |