EDBT 2026 Demo / reviewers in the wild / expert
Yang Wang 0020
dblp:w/YangWang20
· DBLP profile ↗
39ranked-venue papers
2as first author
26since 2021 · last 2026
0000-0002-8903-2388ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 2 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Theory of computation · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Transformers Can Overcome the Curse of Dimensionality: A Theoretical Study from an Approximation PerspectiveabstractThe Transformer model is widely used in various application areas of machine learning, such as natural language processing. This paper investigates the approximation of the Hölder continuous function class $\mathcal{H}_{Q}^{\beta}\left([0,1]^{d\times n},\mathbb{R}^{d\times n}\right)$ by Transformers and constructs several Transformers that can overcome the curse of dimensionality. These Transformers consist of one self-attention layer with one head and the softmax function as the activation function, along with several feedforward layers. For example, to achieve an approximation accuracy of $\epsilon$, if the activation functions of the feedforward layers in the Transformer are ReLU and floor, only $\mathcal{O}\left(\log\frac{1}{\epsilon}\right)$ layers of feedforward layers are needed, with widths of these layers not exceeding $\mathcal{O}\left(\frac{1}{\epsilon^{2/\beta}}\log\frac{1}{\epsilon}\right)$. If other activation functions are allowed in the feedforward layers, the width of the feedforward layers can be further reduced to a constant. These results demonstrate that Transformers have a strong expressive capability. The construction in this paper is based on the Kolmogorov-Arnold Superposition Theorem and does not require the concept of contextual mapping, hence our proof is more intuitively clear compared to previous Transformer approximation works. Additionally, the translation technique proposed in this paper helps to apply the previous approximation results of feedforward neural networks to Transformer research. Yuling Jiao, Yanming Lai, Yang Wang 0020, Bokai Yan |
J. Mach. Learn. Res. | 3 |
| 2025 | Chain of Attack: On the Robustness of Vision-Language Models Against Transfer-Based Adversarial AttacksabstractPre-trained vision-language models (VLMs) have showcased remarkable performance in image and natural language understanding, such as image captioning and response generation. As the practical applications of VLMs become increasingly widespread, their potential safety and robustness issues raise concerns that adversaries may evade the system and cause these models to generate toxic content through malicious attacks. Therefore, evaluating the robustness of open-source VLMs against adversarial attacks has garnered growing attention, with transfer-based attacks as a representative black-box attacking strategy. However, most existing transfer-based attacks neglect the importance of the semantic correlations between vision and text modalities, leading to sub-optimal adversarial example generation and attack performance. To address this issue, we present Chain of Attack (CoA)1, which iteratively enhances the generation of adversarial examples based on the multi-modal semantic update using a series of intermediate attacking steps, achieving superior adversarial transferability and efficiency. A unified attack success rate computing method is further proposed for automatic evasion evaluation. Extensive experiments conducted under the most realistic and high-stakes scenario, demonstrate that our attacking strategy is able to effectively mislead models to generate targeted responses using only black-box attacks without any knowledge of the victim models. The comprehensive robustness evaluation in our paper provides insight into the vulnerabilities of VLMs and offers a reference for the safety considerations of future model developments. Yequan Bie, Jianda Mao, Yangqiu Song, Yang Wang 0020, Hao Chen 0103, Kani Chen |
CVPR | 5 |
| 2025 | Can LLMs Solve Longer Math Word Problems Better?abstractMath Word Problems (MWPs) play a vital role in assessing the capabilities of Large Language Models (LLMs), yet current research primarily focuses on questions with concise contexts. The impact of longer contexts on mathematical reasoning remains under-explored. This study pioneers the investigation of Context Length Generalizability (CoLeG), which refers to the ability of LLMs to solve MWPs with extended narratives. We introduce Extended Grade-School Math (E-GSM), a collection of MWPs featuring lengthy narratives, and propose two novel metrics to evaluate the efficacy and resilience of LLMs in tackling these problems. Our analysis of existing zero-shot prompting techniques with proprietary LLMs along with open-source LLMs reveals a general deficiency in CoLeG. To alleviate these issues, we propose tailored approaches for different categories of LLMs. For proprietary LLMs, we introduce a new instructional prompt designed to mitigate the impact of long contexts. For open-source LLMs, we develop a novel auxiliary task for fine-tuning to enhance CoLeG. Our comprehensive results demonstrate the effectiveness of our proposed methods, showing improved performance on E-GSM. Additionally, we conduct an in-depth analysis to differentiate the effects of semantic understanding and reasoning efficacy, showing that our methods improves the latter. We also establish the generalizability of our methods across several other MWP benchmarks. Our findings highlight the limitations of current LLMs and offer practical solutions correspondingly, paving the way for further exploration of model generalizability and training methodologies. Xin Xu 0001, Tong Xiao 0004, Zitong Chao, Zhenya Huang, Can Yang 0002, Yang Wang 0020 |
ICLR | 6 |
| 2025 | UGPhysics: A Comprehensive Benchmark for Undergraduate Physics Reasoning with Large Language ModelsabstractLarge language models (LLMs) have demonstrated remarkable capabilities in solving complex reasoning tasks, particularly in mathematics. However, the domain of physics reasoning presents unique challenges that have received significantly less attention. Existing benchmarks often fall short in evaluating LLMs’ abilities on the breadth and depth of undergraduate-level physics, underscoring the need for a comprehensive evaluation. To fill this gap, we introduce UGPhysics, a large-scale and diverse benchmark specifically designed to evaluate **U**nder**G**raduate-level **Physics** (**UGPhysics**) reasoning with LLMs. UGPhysics includes 5,520 undergraduate-level physics problems in both English and Chinese across 13 subjects with seven different answer types and four distinct physics reasoning skills, all rigorously screened for data leakage. Additionally, we develop a Model-Assistant Rule-based Judgment (**MARJ**) pipeline specifically tailored for assessing physics problems, ensuring accurate evaluation. Our evaluation of 31 leading LLMs shows that the highest overall accuracy, 49.8% (achieved by OpenAI-o1-mini), emphasizes the need for models with stronger physics reasoning skills, beyond math abilities. We hope UGPhysics, along with MARJ, will drive future advancements in AI for physics reasoning. Codes and data are available at \href{https://github.com/YangLabHKUST/UGPhysics}{https://github.com/YangLabHKUST/UGPhysics}. Xin Xu 0001, Qiyun Xu, Tong Xiao 0004, Tianhao Chen, Shizhe Diao, Can Yang 0002, Yang Wang 0020 |
ICML | 9 |
| 2025 | GPAS: Accelerating Convergence of LLM Pretraining via Gradient-Preserving Activation ScalingabstractModern Large Language Models, such as the LLaMA, Qwen and DeepSeek series, predominantly adopt the Pre-LayerNorm (Pre-LN) Transformer architecture. While being stable during pretraining and scalable to large model sizes, Pre-LN suffers from an exponential growth in activation variance across layers, causing the shortcut to dominate over sub-layer outputs in the residual connection and limiting the learning capacity of deeper layers. To mitigate this issue, we propose Gradient-Preserving Activation Scaling (GPAS), a simple technique that can be used in combination with existing approaches. GPAS works by scaling down the intermediate activations while keeping their gradients unchanged. This leaves information in the activations intact, and avoids the gradient vanishing problem associated with gradient downscaling. Extensive experiments across various model sizes from 71M to 1B show that GPAS achieves consistent performance gains. Beyond enhancing Pre-LN Transformers, GPAS also shows promise in improving alternative architectures such as Sandwich-LN and DeepNorm, demonstrating its versatility and potential for improving training dynamics in a wide range of settings. Our code is available at https://github.com/dandingsky/GPAS. Tianhao Chen, Xin Xu 0001, Zijing Liu, Xinyuan Song 0002, Ajay Jaiswal, Jishan Hu, Yang Wang 0020, Hao Chen 0103, Shizhe Diao, Shiwei Liu 0003, Lu Yin 0006, Can Yang 0002 |
NeurIPS | 9 |
| 2025 | SwitchLingua: The First Large-Scale Multilingual and Multi-Ethnic Code-Switching DatasetabstractCode-switching (CS) is the alternating use of two or more languages within a conversation or utterance, often influenced by social context and speaker identity. This linguistic phenomenon poses challenges for Automatic Speech Recognition (ASR) systems, which are typically designed for a single language and struggle to handle multilingual inputs. The growing global demand for multilingual applications, including Code-Switching ASR (CSASR), Text-to-Speech (TTS), and Cross-Lingual Information Retrieval (CLIR), highlights the inadequacy of existing monolingual datasets. Although some code-switching datasets exist, most are limited to bilingual mixing within homogeneous ethnic groups, leaving a critical need for a large-scale, diverse benchmark akin to ImageNet in computer vision. To bridge this gap, we introduce \textbf{LinguaMaster}, a multi-agent collaboration framework specifically designed for efficient and scalable multilingual data synthesis. Leveraging this framework, we curate \textbf{SwitchLingua}, the first large-scale multilingual and multi-ethnic code-switching dataset, including: (1) 420K CS textual samples across 12 languages, and (2) over 80 hours of audio recordings from 174 speakers representing 18 countries/regions and 63 racial/ethnic backgrounds, based on the textual data. This dataset captures rich linguistic and cultural diversity, offering a foundational resource for advancing multilingual and multicultural research. Furthermore, to address the issue that existing ASR evaluation metrics lack sensitivity to code-switching scenarios, we propose the \textbf{Semantic-Aware Error Rate (SAER)}, a novel evaluation metric that incorporates semantic information, providing a more accurate and context-aware assessment of system performance. Benchmark experiments on SwitchLingua with state-of-the-art ASR models reveal substantial performance gaps, underscoring the dataset’s utility as a rigorous benchmark for CS capability evaluation. In addition, SwitchLingua aims to encourage further research to promote cultural inclusivity and linguistic diversity in speech technology, fostering equitable progress in the ASR field. LinguaMaster (Code): github.com/Shelton1013/SwitchLingua, SwitchLingua (Data): https://huggingface.co/datasets/Shelton1013/SwitchLinguatext, https://huggingface.co/datasets/Shelton1013/SwitchLinguaaudio Xingyuan Liu, Yequan Bie, Tsz Wai Chan, Yangqiu Song, Yang Wang 0020, Hao Chen 0103, Kani Chen |
NeurIPS | 6 |
| 2025 | The Jade Gateway to Trust: Exploring How Socio-Cultural Perspectives Shape Trust Within Chinese NFT CommunitiesabstractToday's world is witnessing an unparalleled rate of technological transformation. The emergence of non-fungible tokens (NFTs) has transformed how we handle digital assets and value. These tokens have captured the interest of scholars and businesspeople alike. However, NFTs have recently seen a sharp decline in popularity. While cryptocurrency volatility and monetary policies greatly influenced NFT market trends, the community aspects of NFT projects--particularly trust-based interactions--also play a crucial role in NFT adoption and sustainability. From a social computing perspective, understanding these trust dynamics offers valuable insights for the development of both the NFT ecosystem and the broader digital economy. China presents a compelling context for examining these dynamics, offering a unique intersection of technological innovation and traditional cultural values. Through an in-depth qualitative study of Chinese NFT communities, we examine how socio-cultural factors influence trust formation and development. We analyzed discussions from eight prominent WeChat groups dedicated to NFTs and conducted 21 semi-structured interviews with three types of NFT community members. We found that trust in Chinese NFT communities is significantly molded by local cultural values. To be precise, Confucian virtues, such as benevolence, propriety , and integrity , play a crucial role in shaping these trust relationships. Our research identifies three critical trust dimensions in China's NFT market: (1) technological , (2) institutional , and (3) social . We examined the challenges in cultivating each dimension. Based on these insights, we developed tailored trust-building guidelines for Chinese NFT stakeholders. These guidelines address trust issues that factor into NFT's declining popularity and could offer valuable strategies for CSCW researchers, developers, and designers aiming to enhance trust in global NFT communities. Our research urges CSCW scholars to take into account the unique socio-cultural contexts when developing trust-enhancing strategies for digital innovations and online interactions. Yifan Cao 0001, Reza Hadi Mogavi, Meng Xia 0002, Leo Yu-Ho Lo, Xiaoqing Zhang 0018, Mei-Jia Lou, Lennart E. Nacke, Yang Wang 0020, Huamin Qu |
Proc. ACM Hum. Comput. Interact. | 8 |
| 2025 | Wasserstein Graph Neural Networks for Graphs With Missing AttributesabstractMissing node attributes pose a common problem in real-world graphs, impacting the performance of graph neural networks' representation learning. Existing GNNs often struggle to effectively leverage incomplete attribute information, as they are not specifically designed for graphs with missing attributes. To address this issue, we propose a novel node representation learning framework called Wasserstein Graph Neural Network (WGNN). Our approach aims to maximize the utility of limited observed attribute information and account for uncertainty caused by missing values. We achieve this by representing nodes as low-dimensional distributions obtained through attribute matrix decomposition. Additionally, we enhance representation expressiveness by introducing a unique message-passing schema that aggregates distributional information from neighboring nodes in the Wasserstein space. We evaluate the performance of WGNN in node classification tasks using both synthetic and real-world datasets under two missing-attribute scenarios. Moreover, we demonstrate the applicability of WGNN in recovering missing values and tackling matrix completion problems, specifically in graphs involving users and items. Experimental results on both tasks convincingly demonstrate the superiority of our proposed method. Tengfei Ma 0001, Yangqiu Song, Yang Wang 0020 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | A Contrast-Saturation Adaptive Model for Low-Light Image EnhancementabstractAbstract. Previous Retinex-based methods simultaneously estimate illumination and reflectance, complicating model design and potentially impacting image enhancement due to flawed prior assumptions. In this paper, we explore the physical basis of low-light image enhancement, focusing on contrast, saturation, and brightness, and propose an adaptive contrast-saturation (ConSat) model. Our ConSat innovation simplifies model design by focusing solely on a brightness enhancement function, allowing for precise image brightness control while keeping colors natural and realistic. We design a new saturation metric that assesses pixel dispersion from the mean of RGB channels, accurately depicting exposure variations across the image. On this basis, we propose a smart contrast stretching function that adapts contrast adjustments to enhance images under varying light. It boosts contrast in dark, low-saturation areas to clarify details and textures, while curbing it in bright, high-saturation areas to prevent overexposure and color distortion. Finally, we employ a pretrained convolutional neural network (CNN)-based denoiser to achieve satisfactory visual appeal. Numerical experiments show that our ConSat is superior to other state-of-the-art methods in brightness improvement, noise removal, and artifact elimination. Qianting Ma, Yang Wang 0020, Tieyong Zeng |
SIAM J. Imaging Sci. | 2 |
| 2025 | Error Analysis of Three-Layer Neural Network Trained With PGD for Deep Ritz MethodabstractMachine learning is a rapidly advancing field with diverse applications across various domains. One prominent area of research is the utilization of deep learning techniques for solving partial differential equations(PDEs). In this work, we specifically focus on employing a three-layer tanh neural network within the framework of the deep Ritz method(DRM) to solve second-order elliptic equations with three different types of boundary conditions. We perform projected gradient descent(PDG) to train the three-layer network and we establish its global convergence. To the best of our knowledge, we are the first to provide a comprehensive error analysis of using overparameterized networks to solve PDE problems, as our analysis simultaneously includes estimates for approximation error, generalization error, and optimization error. We present error bound in terms of the sample size n and our work provides guidance on how to set the network depth, width, step size, and number of iterations for the projected gradient descent algorithm. Importantly, our assumptions in this work are classical and we do not require any additional assumptions on the solution of the equation. This ensures the broad applicability and generality of our results. Yuling Jiao, Yanming Lai, Yang Wang 0020 |
IEEE Trans. Inf. Theory | 3 |
| 2025 | NFTracer: Tracing NFT Impact Dynamics in Transaction-Flow Substitutive Systems With Visual AnalyticsabstractImpact dynamics are crucial for estimating the growth patterns of NFT projects by tracking the diffusion and decay of their relative appeal among stakeholders. Machine learning methods for impact dynamics analysis are incomprehensible and rigid in terms of their interpretability and transparency, whilst stakeholders require interactive tools for informed decision-making. Nevertheless, developing such a tool is challenging due to the substantial, heterogeneous NFT transaction data and the requirements for flexible, customized interactions. To this end, we integrate intuitive visualizations to unveil the impact dynamics of NFT projects. We first conduct a formative study and summarize analysis criteria, including substitution mechanisms, impact attributes, and design requirements from stakeholders. Next, we propose the Minimal Substitution Model to simulate substitutive systems of NFT projects that can be feasibly represented as node-link graphs. Particularly, we utilize attribute-aware techniques to embed the project status and stakeholder behaviors in the layout design. Accordingly, we develop a multi-view visual analytics system, namely NFTracer, allowing interactive analysis of impact dynamics in NFT transactions. We demonstrate the informativeness, effectiveness, and usability of NFTracer by performing two case studies with domain experts and one user study with stakeholders. The studies suggest that NFT projects featuring a higher degree of similarity are more likely to substitute each other. The impact of NFT projects within substitutive systems is contingent upon the degree of stakeholders' influx and projects' freshness. Yifan Cao 0001, Lue Shen, Kani Chen, Yang Wang 0020, Wei Zeng 0004, Huamin Qu |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2024 | RL in Markov Games with Independent Function Approximation: Improved Sample Complexity Bound under the Local Access ModelabstractEfficiently learning equilibria with large state and action spaces in general-sum Markov games while overcoming the curse of multi-agency is a challenging problem. Recent works have attempted to solve this problem by employing independent linear function classes to approximate the marginal $Q$-value for each agent. However, existing sample complexity bounds under such a framework have a suboptimal dependency on the desired accuracy $\varepsilon$ or the action space. In this work, we introduce a new algorithm, Lin-Confident-FTRL, for learning coarse correlated equilibria (CCE) with local access to the simulator, i.e., one can interact with the underlying environment on the visited states. Up to a logarithmic dependence on the size of the state space, Lin-Confident-FTRL learns $\epsilon$-CCE with a provable optimal accuracy bound $O(\epsilon^{-2})$ and gets rids of the linear dependency on the action space, while scaling polynomially with relevant problem parameters (such as the number of agents and time horizon). Moreover, our analysis of Linear-Confident-FTRL generalizes the virtual policy iteration technique in the single-agent local planning literature, which yields a new computationally efficient algorithm with a tighter sample complexity bound when assuming random access to the simulator. Junyi Fan, Jialin Zeng, Jian-Feng Cai 0001, Yang Wang 0020, Jiheng Zhang |
AISTATS | 5 |
| 2024 | On the Convergence of Projected Bures-Wasserstein Gradient Descent under Euclidean Strong ConvexityabstractThe Bures-Wasserstein (BW) gradient descent method has gained considerable attention in various domains, including Gaussian barycenter, matrix recovery and variational inference problems, due to its alignment with the Wasserstein geometry of normal distributions. Despite its popularity, existing convergence analysis are often contingent upon specific loss functions, and the exploration of constrained settings within this framework remains limited. In this work, we make an attempt to bridge this gap by providing a general convergence rate guarantee for BW gradient descent when the Euclidean strong convexity of the loss and the constraints is assumed. In an effort to advance practical implementations, we also derive a closed-form solution for the projection onto BW distance-constrained sets, which enables the fast implementation of projected BW gradient descent for problems that arise in the constrained barycenter and distributionally robust optimization literature. Experimental results demonstrate significant improvements in computational efficiency and convergence speed, underscoring the efficacy of our method in practical scenarios. Junyi Fan, Zijian Liu 0003, Jian-Feng Cai 0001, Yang Wang 0020, Zhengyuan Zhou |
ICML | 5 |
| 2024 | PoeticAR: Reviving Traditional Poetry of the Heritage Site of Jichang Garden via Augmented RealityabstractAs a famed Chinese classical garden, the Jichang Garden was a constant inspiration to many poets in its hundreds of years’ history, who composed a rich body of poems—a valuable intangible cultural heritage. While tourists tend to pay attention to tangible natural scenery and historical architectures, they often neglect intangible cultural heritage—poems. We interviewed 23 tourists and found that augmented reality (AR) was viable for tourists to enjoy the physical scenery and the poetry simultaneously. We developed an initial prototype of PoeticAR, which presents poems based on physical scenery to enhance tourists’ cultural and aesthetic experience. We further revised the prototype based on the ideas generated from a workshop with 18 tourists. We conducted a between-subject user study with 30 tourists to compare PoeticAR with Video. Results showed that PoeticAR significantly motivated tourists’ interest in poems, enhanced the cultural and aesthetic tour experience in Jichang Garden, and increased awareness of Intangible Cultural Heritage of Cultural Heritage sites. Yifan Cao 0001, Lingyi Feng, Dongting Fu, Linping Yuan, Huamin Qu, Yang Wang 0020, Mingming Fan 0001 |
Int. J. Hum. Comput. Interact. | 7 |
| 2023 | Optimal Contextual Bandits with Knapsacks under Realizability via Regression OraclesabstractWe study the stochastic contextual bandit with knapsacks (CBwK) problem, where each action, taken upon a context, not only leads to a random reward but also costs a random resource consumption in a vector form. The challenge is to maximize the total reward without violating the budget for each resource. We study this problem under a general realizability setting where the expected reward and expected cost are functions of contexts and actions in some given general function classes $\mathcal{F}$ and $\mathcal{G}$, respectively. Existing works on CBwK are restricted to the linear function class since they use UCB-type algorithms, which heavily rely on the linear form and thus are difficult to extend to general function classes. Motivated by online regression oracles that have been successfully applied to contextual bandits, we propose the first universal and optimal algorithmic framework for CBwK by reducing it to online regression. We also establish the lower regret bound to show the optimality of our algorithm for a variety of function classes. Jialin Zeng, Yang Wang 0020, Jiheng Zhang |
AISTATS | 3 |
| 2023 | NFTeller: Dual-centric Visual Analytics for Assessing Market Performance of NFT CollectiblesabstractNon-fungible tokens (NFTs) have recently gained widespread popularity as an alternative investment. However, the lack of assessment criteria has caused intense volatility in NFT marketplaces. Identifying attributes impacting the market performance of NFT collectibles is crucial but challenging due to the massive amount of heterogeneous and multi-modal data in NFT transactions, e.g., social media texts, numerical trading data, and images. To address this challenge, we introduce an interactive dual-centric visual analytics system, NFTeller, to facilitate users’ analysis. First, we collaborate with five domain experts to distill static and dynamic impact attributes and collect relevant data. Next, we derive six analysis tasks and develop NFTeller to present the evolution of NFT transactions and correlate NFTs’ market performance with impact attributes. Notably, we create an augmented chord diagram with a radial stacked bar chart to explore intersections between NFT collection projects and whale accounts. Finally, we conduct three case studies and interview domain experts to evaluate the effectiveness and usability of this system. As such, we gain in-depth insights into assessing NFT collectibles and detecting opportune moments for investment. Yifan Cao 0001, Meng Xia 0002, Kento Shigyo, Furui Cheng, Qianhang Yu, Xingxing Yang 0005, Yang Wang 0020, Wei Zeng 0004, Huamin Qu |
VINCI | 7 |
| 2023 | Retinex-Based Variational Framework for Low-Light Image Enhancement and DenoisingabstractLow-light image enhancement is an important task in the domain of computer vision. Images taken under insufficient lighting conditions manifest low visibility and unknown noises which disrupt image contents and pose considerable challenges for low-light image enhancement. Most of Retinex-based methods usually attempt to design different priors on the gradient of both illumination and reflectance. However, noises can be involved in the Retinex-based models. To address the problem, we explore the problem of low-light image restoration through joint contrast enhancement and denoising. We propose a Retinex-based variational model for low-light image enhancement that effectively generates a noise-free image, yet proves to generalize well to diverse light-conditions. First, we present a simple constraint on the fidelity term between the fractional derivative of an observed image and the fractional derivative of the recomposed one which is the product of the reflectance and illumination. This strategy aims to model spatial consistency to preserve natural variation. Second, we introduce a weighted regularization term for the reflectance that can remove noise with a adaptive texture map. We evaluate our proposed approach using three challenging datasets: NPE, LOL and GladNet. Extensive experiments demonstrate that our proposed method outperforms other competing methods in terms of visual quality and quantitative comparisons. Qianting Ma, Yang Wang 0020, Tieyong Zeng |
IEEE Trans. Multim. | 2 |
| 2022 | Variance-Reduced Randomized Kaczmarz Algorithm In Xfel Single-Particle Imaging Phase RetrievalabstractIn this paper, we propose the Variance Reduced Randomized Kaczmarz (VR-RK) algorithm for XFEL signal particle imaging phase retrieval. The VR-RK algorithm is inspired by the randomized Kaczmarz algorithm and the variance reduction in stochastic gradient methods. The formulations of the VR-RK algorithm under the L1and L2constraints are also presented. Numerical simulations demonstrate that the VR-RK method has a faster convergence rate compared with the randomized Kaczmarz method. Tests on the synthetic signal particle imaging data and the PR772 XFEL real imaging data show that the VR-RK algorithm can recover information with higher accuracy. It is useful for biological data processing. Yin Xian, Haiguang Liu, Xue-Cheng Tai, Yang Wang 0020 |
ICIP | 4 |
| 2022 | Private Streaming SCO in ℓp geometry with Applications in High Dimensional Online Decision Making
Zhicong Liang, Yang Wang 0020, Yuan Yao 0011, Jiheng Zhang |
ICML | 4 |
| 2022 | An Error Analysis of Generative Adversarial Networks for Learning DistributionsabstractThis paper studies how well generative adversarial networks (GANs) learn probability distributions from finite samples. Our main results establish the convergence rates of GANs under a collection of integral probability metrics defined through Hölder classes, including the Wasserstein distance as a special case. We also show that GANs are able to adaptively learn data distributions with low-dimensional structures or have Hölder densities, when the network architectures are chosen properly. In particular, for distributions concentrated around a low-dimensional set, we show that the learning rates of GANs do not depend on the high ambient dimension, but on the lower intrinsic dimension. Our analysis is based on a new oracle inequality decomposing the estimation error into the generator and discriminator approximation error and the statistical error, which may be of independent interest. Jian Huang 0003, Yuling Jiao, Zhen Li 0074, Shiao Liu, Yang Wang 0020, Yunfei Yang 0002 |
J. Mach. Learn. Res. | 5 |
| 2022 | On the capacity of deep generative networks for approximating distributions
Yunfei Yang 0002, Zhen Li 0074, Yang Wang 0020 |
Neural Networks | 3 |
| 2022 | Approximation in shift-invariant spaces with deep ReLU neural networks
Yunfei Yang 0002, Zhen Li 0074, Yang Wang 0020 |
Neural Networks | 3 |
| 2022 | Linear Convergence of Randomized Kaczmarz Method for Solving Complex-Valued Phaseless EquationsabstractA randomized Kaczmarz method was recently proposed for phase retrieval, which has been shown numerically to exhibit empirical performance over other state-of-the-art phase retrieval algorithms both in terms of the sampling complexity and computation time. While the rate of convergence has been well studied in the real case where the signals and measurement vectors are all real-valued, there is no guarantee for the convergence in the complex case. In fact, the linear convergence of the randomized Kaczmarz method for phase retrieval in the complex setting is left as a conjecture by Tan and Vershynin [ Inf. Inference, 8 (2019), pp. 97--123]. In this paper, we provide the first theoretical guarantees for it. We show that for random measurements ${a}_j \in \mathbb{C}^n,\, j=1,\ldots,m,$ which are drawn independently and uniformly from the complex unit sphere, or equivalently are independent complex Gaussian random vectors, when $m\ge Cn$ for some universal positive constant $C$, the randomized Kaczmarz scheme with a good initialization converges linearly to the target solution (up to a global phase) in expectation with high probability. This gives a positive answer to that conjecture. Meng Huang 0002, Yang Wang 0020 |
SIAM J. Imaging Sci. | 2 |
| 2021 | Deep Generative Learning via Schrödinger BridgeabstractWe propose to learn a generative model via entropy interpolation with a Schr{ö}dinger Bridge. The generative learning task can be formulated as interpolating between a reference distribution and a target distribution based on the Kullback-Leibler divergence. At the population level, this entropy interpolation is characterized via an SDE on [0,1] with a time-varying drift term. At the sample level, we derive our Schr{ö}dinger Bridge algorithm by plugging the drift term estimated by a deep score estimator and a deep density ratio estimator into the Euler-Maruyama method. Under some mild smoothness assumptions of the target distribution, we prove the consistency of both the score estimator and the density ratio estimator, and then establish the consistency of the proposed Schr{ö}dinger Bridge approach. Our theoretical results guarantee that the distribution learned by our approach converges to the target distribution. Experimental results on multimodal synthetic data and benchmark data support our theoretical findings and indicate that the generative model via Schr{ö}dinger Bridge is comparable with state-of-the-art GANs, suggesting a new formulation of generative learning. We demonstrate its usefulness in image interpolation and image inpainting. Gefei Wang, Yuling Jiao, Yang Wang 0020, Can Yang 0002 |
ICML | 4 |
| 2021 | Generalized Linear Bandits with Local Differential PrivacyabstractContextual bandit algorithms are useful in personalized online decision-making. However, many applications such as personalized medicine and online advertising require the utilization of individual-specific information for effective learning, while user's data should remain private from the server due to privacy concerns. This motivates the introduction of local differential privacy (LDP), a stringent notion in privacy, to contextual bandits. In this paper, we design LDP algorithms for stochastic generalized linear bandits to achieve the same regret bound as in non-privacy settings. Our main idea is to develop a stochastic gradient-based estimator and update mechanism to ensure LDP. We then exploit the flexibility of stochastic gradient descent (SGD), whose theoretical guarantee for bandit problems is rarely explored, in dealing with generalized linear bandits. We also develop an estimator and update mechanism based on Ordinary Least Square (OLS) for linear bandits. Finally, we conduct experiments with both simulation and real-world datasets to demonstrate the consistently superb performance of our algorithms under LDP constraints with reasonably small parameters $(\varepsilon, \delta)$ to ensure strong privacy protection. Yang Wang 0020, Jiheng Zhang |
NeurIPS | 3 |
| 2021 | Non-asymptotic Error Bounds for Bidirectional GANsabstractWe derive nearly sharp bounds for the bidirectional GAN (BiGAN) estimation error under the Dudley distance between the latent joint distribution and the data joint distribution with appropriately specified architecture of the neural networks used in the model. To the best of our knowledge, this is the first theoretical guarantee for the bidirectional GAN learning approach. An appealing feature of our results is that they do not assume the reference and the data distributions to have the same dimensions or these distributions to have bounded support. These assumptions are commonly assumed in the existing convergence analysis of the unidirectional GANs but may not be satisfied in practice. Our results are also applicable to the Wasserstein bidirectional GAN if the target distribution is assumed to have a bounded support. To prove these results, we construct neural network functions that push forward an empirical distribution to another arbitrary empirical distribution on a possibly different-dimensional space. We also develop a novel decomposition of the integral probability metric for the error analysis of bidirectional GANs. These basic theoretical results are of independent interest and can be applied to other related learning problems. Shiao Liu, Yunfei Yang 0002, Jian Huang 0003, Yuling Jiao, Yang Wang 0020 |
NeurIPS | 5 |
| 2020 | Learning Hybrid Representations for Automatic 3D Vessel Centerline Extraction
Jiafa He, Chengwei Pan, Can Yang 0002, Ming Zhang 0004, Yang Wang 0020, Xiaowei Zhou 0001, Yizhou Yu |
MICCAI (6) | 5 |
| 2020 | Variational Single Image Dehazing for Enhanced VisualizationabstractIn this paper, we investigate the challenging task of removing haze from a single natural image. The analysis on the haze formation model shows that the atmospheric veil has much less relevance to chrominance than luminance, which motivates us to neglect the haze in the chrominance channel and concentrate on the luminance channel in the dehazing process. Besides, the experimental study illustrates that the YUV color space is most suitable for image dehazing. Accordingly, a variational model is proposed in the Y channel of the YUV color space by combining the reformulation of the haze model and the two effective priors. As we mainly focus on the Y channel, most of the chrominance information of the image is preserved after dehazing. The numerical procedure based on the alternating direction method of multipliers (ADMM) scheme is presented to obtain the optimal solution. Extensive experimental results on real-world hazy images and synthetic dataset demonstrate clearly that our method can unveil the details and recover vivid color information, which is competitive among many existing dehazing algorithms. Further experiments show that our model also can be applied for image enhancement. Faming Fang, Tingting Wang 0007, Yang Wang 0020, Tieyong Zeng, Guixu Zhang |
IEEE Trans. Multim. | 3 |
| 2019 | Convex Shape Prior for Multi-Object Segmentation Using a Single Level Set FunctionabstractMany objects in real world have convex shapes. It is a difficult task to have representations for convex shapes with good and fast numerical solutions. This paper proposes a method to incorporate convex shape prior for multi-object segmentation using level set method. The relationship between the convexity of the segmented objects and the signed distance function corresponding to their union is analyzed theoretically. This result is combined with Gaussian mixture method for the multiple objects segmentation with convexity shape prior. Alternating direction method of multiplier (ADMM) is adopted to solve the proposed model. Special boundary conditions are also imposed to obtain efficient algorithms for 4th order partial differential equations in one step of ADMM algorithm. In addition, our method only needs one level set function regardless of the number of objects. So the increase in the number of objects does not result in the increase of model and algorithm complexity. Various numerical experiments are illustrated to show the performance and advantages of the proposed method. Shousheng Luo, Xue-Cheng Tai, Limei Huo, Yang Wang 0020, Roland Glowinski |
ICCV | 4 |
| 2019 | Deep Generative Learning via Variational Gradient FlowabstractWe propose a framework to learn deep generative models via \textbf{V}ariational \textbf{Gr}adient Fl\textbf{ow} (VGrow) on probability spaces. The evolving distribution that asymptotically converges to the target distribution is governed by a vector field, which is the negative gradient of the first variation of the $f$-divergence between them. We prove that the evolving distribution coincides with the pushforward distribution through the infinitesimal time composition of residual maps that are perturbations of the identity map along the vector field. The vector field depends on the density ratio of the pushforward distribution and the target distribution, which can be consistently learned from a binary classification problem. Connections of our proposed VGrow method with other popular methods, such as VAE, GAN and flow-based methods, have been established in this framework, gaining new insights of deep generative learning. We also evaluated several commonly used divergences, including Kullback-Leibler, Jensen-Shannon, Jeffreys divergences as well as our newly discovered “logD” divergence which serves as the objective function of the logD-trick GAN. Experimental results on benchmark datasets demonstrate that VGrow can generate high-fidelity images in a stable and efficient manner, achieving competitive performance with state-of-the-art GANs. Yuan Gao 0044, Yuling Jiao, Yang Wang 0020, Yao Wang 0003, Can Yang 0002, Shunkang Zhang |
ICML | 3 |
| 2018 | Scattering transform and sparse linear classifiers for art authentication
Roberto F. Leonarduzzi, Yang Wang 0020 |
Signal Process. | 3 |
| 2013 | Bayesian Learning of Sparse Multiscale Image RepresentationsabstractMultiscale representations of images have become a standard tool in image analysis. Such representations offer a number of advantages over fixed-scale methods, including the potential for improved performance in denoising, compression, and the ability to represent distinct but complementary information that exists at various scales. A variety of multiresolution transforms exist, including both orthogonal decompositions such as wavelets as well as nonorthogonal, overcomplete representations. Recently, techniques for finding adaptive, sparse representations have yielded state-of-the-art results when applied to traditional image processing problems. Attempts at developing multiscale versions of these so-called dictionary learning models have yielded modest but encouraging results. However, none of these techniques has sought to combine a rigorous statistical formulation of the multiscale dictionary learning problem and the ability to share atoms across scales. We present a model for multiscale dictionary learning that overcomes some of the drawbacks of previous approaches by first decomposing an input into a pyramid of distinct frequency bands using a recursive filtering scheme, after which we perform dictionary learning and sparse coding on the individual levels of the resulting pyramid. The associated image model allows us to use a single set of adapted dictionary atoms that is shared--and learned--across all scales in the model. The underlying statistical model of our proposed method is fully Bayesian and allows for efficient inference of parameters, including the level of additive noise for denoising applications. We apply the proposed model to several common image processing problems including non-Gaussian and nonstationary denoising of real-world color images. James Michael Hughes, Daniel N. Rockmore, Yang Wang 0020 |
IEEE Trans. Image Process. | 3 |
| 2012 | Empirical Mode Decomposition Analysis for Visual StylometryabstractIn this paper, we show how the tools of empirical mode decomposition (EMD) analysis can be applied to the problem of “visual stylometry,” generally defined as the development of quantitative tools for the measurement and comparisons of individual style in the visual arts. In particular, we introduce a new form of EMD analysis for images and show that it is possible to use its output as the basis for the construction of effective support vector machine (SVM)-based stylometric classifiers. We present the methodology and then test it on collections of two sets of digital captures of drawings: a set of authentic and well-known imitations of works attributed to the great Flemish artist Pieter Bruegel the Elder (1525-1569) and a set of works attributed to Dutch master Rembrandt van Rijn (1606-1669) and his pupils. Our positive results indicate that EMD-based methods may hold promise generally as a technique for visual stylometry. James Michael Hughes, Dong Mao, Daniel N. Rockmore, Yang Wang 0020, Qiang Wu 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2010 | The golden ratio encoderabstractThis paper proposes a novel Nyquist-rate analog-to-digital (A/D) conversion algorithm which achieves exponential accuracy in the bit-rate despite using imperfect components. The proposed algorithm is based on a robust implementation of a beta-encoder with β = φ = (1 + √5)/2, the golden ratio. It was previously shown that beta-encoders can be implemented in such a way that their exponential accuracy is robust against threshold offsets in the quantizer element. This paper extends this result by allowing for imperfect analog multipliers with imprecise gain values as well. Furthermore, a formal computational model for algorithmic encoders and a general test bed for evaluating their robustness is proposed. Ingrid Daubechies, C. Sinan Güntürk, Yang Wang 0020, Özgür Yilmaz |
IEEE Trans. Inf. Theory | 3 |
| 2003 | Substitution Delone Sets
Jeffrey C. Lagarias, Yang Wang 0020 |
Discret. Comput. Geom. | 2 |
| 2001 | Disk-Like Self-Affine Tiles in R2
Christoph Bandt, Yang Wang 0020 |
Discret. Comput. Geom. | 2 |
| 1998 | On the use of high-order ambiguity function for multi-component polynomial phase signals
Yang Wang 0020, G. Tong Zhou |
Signal Process. | 1 |
| 1997 | On the use of high order ambiguity function for multicomponent polynomial phase signalsabstractNonstationary signals appear often in real-life applications and many of them can be modeled as polynomial phase signals (PPS). The high-order ambiguity function (HAF) was first introduced to estimate the parameters of single component PPS. But due to its high nonlinearity, HAF has not been widely used for multi-component PPS which appear, for example, in Doppler radar applications when multiple targets are tracked simultaneously. We present a theory in this paper that HAF is virtually additive for multi-component PPS and illustrate our findings with numerical simulations. Yang Wang 0020, G. Tong Zhou |
ICASSP | 1 |
| 1997 | Exploring lag diversity in the high-order ambiguity function for polynomial phase signalsabstractHigh-order ambiguity function (HAF) is an effective tool for retrieving coefficients of polynomial phase signals (PPS's). The lag choice is dictated by conflicting requirements: a large lag improves estimation accuracy but drastically limits the range of the parameters that can be estimated. By using two (large) coprime lags and solving linear Diophantine equations using the Euclidean algorithm, we are able to recover the PPS coefficients from aliased peak positions without compromising the dynamic range and the estimation accuracy. Separating components of a multicomponent PPS whose phase polynomials have very similar leading coefficients has been a challenging task, but can now be tackled easily with the two-lag approach. Numerical examples are presented to illustrate the effectiveness of our method. G. Tong Zhou, Yang Wang 0020 |
IEEE Signal Process. Lett. | 2 |