EDBT 2026 Demo / reviewers in the wild / expert
Miao Lu
dblp:09/1168
· DBLP profile ↗
30ranked-venue papers
8as first author
18since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 4 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 3 since 2021Computer networks · 3 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Databases, data management, data science and information retrieval · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond the Context Window: Scaling Agentic RL via End-to-end Optimized Context CompressionabstractWe study reinforcement learning (RL) finetuning of large language model (LLM) agents for long-horizon multi-turn tool use, where context length quickly becomes a fundamental bottleneck.Existing multi-turn RL pipelines suffer from degraded instruction following, excessive rollout costs, and most importantly, strict context limits.In this work, to address these challenges, we introduce summarization-based context management to training.In specific, it periodically compresses the tool using history by LLM-generated summaries that retain taskrelevant information to keep a compact context while enabling the agent to scale beyond the fixed context window.Building on this formulation, we derive a policy gradient representation that seamlessly enables standard LLM RL infrastructures to optimize both tool-use behaviors as well as summarization strategies in an end-to-end fashion.We instantiate this framework with SUmmarization augmented Policy Optimization (SUPO), an LLM RL algorithm that enables long-horizon training beyond a fixed context limit.Experiments on interactive function calling and searching tasks demonstrate that SUPO significantly improves the success rate while maintaining the same or even lower working context length compared to baselines.We also demonstrate that for complex searching tasks SUPO can further improve the evaluation performance when scaling test-time maximum round of summarization beyond that of training time. Miao Lu, Weihua Du, Zhan Ling, Xuesong Yao, Jiecao Chen |
ACL (1) | 1 |
| 2025 | Can Neural Networks Achieve Optimal Computational-statistical Tradeoff? An Analysis on Single-Index ModelabstractIn this work, we tackle the following question: Can neural networks trained with gradient-based methods achieve the optimal statistical-computational tradeoff in learning Gaussian single-index models?
Prior research has shown that any polynomial-time algorithm under the statistical query (SQ) framework requires $\Omega(d^{s^\star/2}\lor d)$ samples, where $s^\star$ is the generative exponent representing the intrinsic difficulty of learning the underlying model.
However, it remains unknown whether neural networks can achieve this sample complexity.
Inspired by prior techniques such as label transformation and landscape smoothing for learning single-index models, we propose a unified gradient-based algorithm for training a two-layer neural network in polynomial time.
Our method is adaptable to a variety of loss and activation functions, covering a broad class of existing approaches.
We show that our algorithm learns a feature representation that strongly aligns with the unknown signal $\theta^\star$, with sample complexity $\tilde O (d^{s^\star/2} \lor d)$, matching the SQ lower bound up to a polylogarithmic factor for all generative exponents $s^\star\geq 1$.
Furthermore, we extend our approach to the setting where $\theta^\star$ is $k$-sparse for $k = o(\sqrt{d})$ by introducing a novel weight perturbation technique that leverages the sparsity structure.
We derive a corresponding SQ lower bound
of order $\tilde\Omega(k^{s^\star})$, matched by our method up to a polylogarithmic factor.
Our framework, especially the weight perturbation technique, is of independent interest, and suggests potential gradient-based solutions to other problems such as sparse tensor PCA. Siyu Chen 0001, Beining Wu, Miao Lu, Zhuoran Yang, Tianhao Wang 0002 |
ICLR | 3 |
| 2025 | Multiobjective Optimization and Decision-Making of Low-Carbon Environmental Parameters in AIoT-Enabled Controlled Environment AgricultureabstractCo-optimizing air temperature (AT) and CO2 concentration (CO2) for AIoT systems in controlled environment agriculture (CEA) has raised sustained attentions to balance energy efficiency and crop productivity. To tackle this issue, we propose a region-based co-optimization strategy that combines crop-environment interaction modeling with a multi-objective optimization algorithm to obtain the optimal ranges of AT and CO2. Experimental data are collected by measuring single-leaf photosynthetic rate (AL) in lettuce cultivated under diverse environmental conditions. AL is used as the primary crop growth indicator and modeled using support vector regression (SVR). Based on the prediction model, novel fitness functions and constraints are established considering biomass accumulation, energy consumption, and AIoT system actions. Subsequently, a hybrid multi-stage method integrating the non-dominated sorting genetic algorithm II (NSGAII) and differential evolutionary (DE) algorithm is employed to determine the non-dominated solution (NDS) sets for different conditions. Finally, the optimal solution within the NDS sets is calculated using the VlseKriterijumska Optimizacija I Kompromisno Resenje (VIKOR) method constrained by curvature feature. The rationalities of the proposed methods are verified by lettuce and cucumber samples. Experimental validation with lettuce cultivation across multiple environments demonstrates significant advantages over conventional threshold-based methods: 18.0% improvement in carbon use efficiency, 40.7% reduction in AIoT system operations, and 94.7% photosynthetic efficiency retention. This research achieves a groundbreaking integration of agronomic models, intelligent algorithms, and AIoT systems, establishing a novel approach for sustainable agricultural management that simultaneously optimizes production efficiency and resource utilization. Miao Lu, Haoling Liu, Yongxia Yang, Huimin Li 0009, Pan Gao 0003, Jin Hu 0007 |
IEEE Internet Things J. | 1 |
| 2025 | Real-Time Nitrogen Regulation via IoT Edge Computing: A Chlorophyll Fluorescence-Driven Framework for Sustainable Plant FactoriesabstractTraditional nitrogen management systems in plant factories, based on static nutrient formulations, cause 30 50% nitrogen wastage and environmental pollution. Fixed threshold parameters fail to meet the dynamic needs of crops, reducing yield and quality, this study proposes an Internet of Things (IoT) closed-loop control system based on a dynamic nitrogen regulation model. By integrating Maximum Information Coefficient (MIC), Analytic Hierarchy Process (AHP), U-chord curvature method, and Technique for Order Preference by Similarity to Ideal Solution (TOPSIS), a multi-parameter dynamic weighting nitrogen regulation interval model was developed. This model overcomes the limitations of traditional methods that rely on fixed formulas or single parameters. The experiments demonstrated that the model can accurately identify the optimal nitrogen concentrations during the seedling stage (5.0 8.0 mmol/L) and the maturity stage (11.0 14.0 mmol/L), reducing nitrogen fertilizer usage by 36.2% 49.7%, while increasing fresh weight by 12% and leaf potassium and phosphorus contents by 4.8% 19.5%. The system employs an edge-cloud collaborative architecture, with the Raspberry Pi 4B serving as the edge node for real-time regulation (response time. 0.5 seconds), supporting remote monitoring and model updates, forming a "perception-decision-execution" closed-loop. This research provides a nitrogen management paradigm for smart agriculture that combines dynamic precision with engineering practicality, which can be extended to vertical farms and urban agricultural scenarios. Zhangtong Sun, Yongxia Yang, Miao Lu, Huimin Li 0009, Jie-Xiao Peng, Shijie Tian, Jin Hu 0007, Pan Gao 0003 |
IEEE Internet Things J. | 3 |
| 2025 | Evaluating Sustainable Development in the Middle and Lower Reaches of the Yellow River Basin Using Multiple Data SourcesabstractThe Yellow River Basin is an important area in China’s overall development, especially the middle and lower reaches of this river which are vital for the development of the entire basin. Therefore, it is critical to scientifically evaluate the sustainable development of the ecological environment and the social economy in these reaches of the Yellow River Basin. In this study, the level of the coordinated and sustainable development of the ecology and social economy was quantitatively analyzed based on the improved remote sensing ecological index (IRSEI) and development space social and economic curve (DSSEC). The analysis process is as follows. First, Landsat 8 data were used to invert the surface wetness index (WET), normalized difference building and soil index (NDBSI), land surface temperature index (LST), and kernel normalized difference vegetation index (kNDVI). Then, the IRSEI was constructed. Combining principal component analysis (PCA) and particle swarm optimization (PSO) random forest algorithm, an ecological quality evaluation model was established to evaluate the ecological environment in the middle and lower reaches of the Yellow River Basin. The dynamic temporal and spatial evolutions of the ecological environment quality and ecological sustainability were determined. Second, the relationship between the development space and two key social economy processes was quantitatively determined using the DSSEC index, in which the coefficients reflect the basic intensity and concentration of socioeconomic activities, enabling the analysis of the region’s socioeconomic sustainability. Third, the level of coordinated and sustainable ecological and socioeconomic development and its influencing factors were studied using a coupling coordination degree model and a geographical detector. The results showed that the IRSEI in the study area gradually increased from 0.392 in 2014 to 0.612 in 2020, and the ecological quality level in most areas was moderate to good. The social and economic development in the middle and lower reaches of the Yellow River Basin had regional urban centers as the core, indicating the need to emphasize the sustainable development of the central area in planning and development strategies to promote coordinated and balanced development. From 2014 to 2020, we identified an overall trend of linear growth, and the degree of coordination gradually transitioned from moderate imbalance to moderate coordination, showing a trend of increasing sustainable development. The degree of water resource development, industrial structure, and level of urbanization are the main factors driving the coordinated and sustainable development of the ecology and social economy in the middle and lower reaches of the Yellow River Basin. Yuefeng Lu, Zhenqi Song, Miao Lu |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | Research on Desertification Monitoring and Vegetation Refinement Extraction Methods Based on the Synergy of Multisource Remote Sensing ImageryabstractDue to over-exploitation by humans and global climate change, desertification has become an increasingly severe issue, seriously threatening the stability of ecosystems and the sustainable development of resources. Therefore, this study focuses on the Hangjin Banner region in Inner Mongolia, using satellite remote sensing and remote aerial vehicles (RAV) remote sensing technology. Through wide-area coverage, long-term monitoring, multiscale analysis, and high-precision interpretation, the study demonstrates the strong synergistic effects of “multiscale interpretation” and “data fusion applications,” systematically carrying out desertification monitoring grading and refined vegetation extraction. First, to address the problem that the information dimension of a single index is insufficient and it is difficult to reflect the development trend of desertification, the normalized difference vegetation index (NDVI)-albedo feature space applicable to the desert environment is inversely performed based on Landsat 8 satellite images from 2009 to 2023. Then, on the basis of the feature space, the desertification difference index (DDI), which realizes the wide-area desertification monitoring grading and spatio-temporal evolution analysis of the study area, and the hue-saturation-lightness greenway enhanced vegetation index (HSLGEVI), which has stronger applicability and stability in desert environments, were constructed based on the HSL color space and the hue tuning algorithm. This index can effectively overcome the limitations of the RGB vegetation index, clearly delineate the canopy edge of desert vegetation, and accurately extract surface meadow vegetation with lower chlorophyll content. To test the effectiveness of the HSLGEVI, the widely used and validated excess green index (EXG), vegetation difference vegetation index (VDVI), modified green-red vegetation index (MGRVI), and red-green–blue vegetation index (RGBVI) were selected for comparison. The results show that the accuracy of HSLGEVI is better than that of other indices, with overall accuracy and${F}1$-score remaining above 90%. It reduces the impact of the RGB color space vegetation index on the accuracy of vegetation extraction, effectively overcoming misclassification and omission issues, and providing a reliable monitoring mechanism for desertification control in the Hangjin Banner area. Zhenqi Song, Yuefeng Lu, Jinhui Yuan, Miao Lu, Dengkuo Sun |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Benign Oscillation of Stochastic Gradient Descent with Large Learning RateabstractIn this work, we theoretically investigate the generalization properties of neural networks (NN) trained by stochastic gradient descent (SGD) with large learning rates. Under such a training regime, our finding is that, the oscillation of the NN weights caused by SGD with large learning rates turns out to be beneficial to the generalization of the NN, potentially improving over the same NN trained by SGD with small learning rates that converges more smoothly. In view of this finding, we call such a phenomenon “benign oscillation”. Our theory towards demystifying such a phenomenon builds upon the feature learning perspective of deep learning. Specifically, we consider a feature-noise data generation model that consists of (i) weak features which have a small $\ell_2$-norm and appear in each data point; (ii) strong features which have a large $\ell_2$-norm but appear only in a certain fraction of all data points; and (iii) noise. We prove that NNs trained by oscillating SGD with a large learning rate can effectively learn the weak features in the presence of those strong features. In contrast, NNs trained by SGD with a small learning rate can only learn the strong features but make little progress in learning the weak features. Consequently, when it comes to the new testing data points that consist of only weak features, the NN trained by oscillating SGD with a large learning rate can still make correct predictions, while the NN trained by SGD with a small learning rate could not. Our theory sheds light on how large learning rate training benefits the generalization of NNs. Experimental results demonstrate our findings on the phenomenon of “benign oscillation”. Miao Lu, Beining Wu, Difan Zou |
ICLR | 1 |
| 2024 | Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial RegularizerabstractAligning generative models with human preference via RLHF typically suffers from overoptimization, where an imperfectly learned reward model can misguide the generative model to output even undesired responses. We investigate this problem in a principled manner by identifying the source of the issue as the distributional shift and uncertainty of human preference in dataset. To mitigate overoptimization, we first propose a theoretical algorithm which optimizes the policy against an adversarially chosen reward model, one that simultaneously minimizes its MLE loss and a reward penalty term. The penalty pessimistically biases the uncertain rewards so as to prevent the policy from choosing actions with spursiouly high proxy rewards, resulting in provable sample efficiency of the algorithm under a partial coverage style condition. Moving from theory to practice, the proposed algorithm further enjoys an equivalent but surprisingly easy to implement form. With a clever usage of the equivalence between reward models and the corresponding optimal policy, the algorithm features a simple objective that combines (i) a preference optimization loss that directly aligns the policy with human preference, and (ii) a supervised learning loss which explicitly imitates the policy with a baseline distribution. In the context of aligning large language models (LLM), this objective fuses the direct preference optimization (DPO) loss with the supervised fune-tuning (SFT) loss to help mitigate the overoptimization towards undesired responses, for which we name the algorithm Regularized Preference Optimization (RPO).
Experiments of aligning LLMs demonstrate the improved performance of our method when compared with DPO baselines.
Our work sheds light on the interplay between preference optimization and SFT in tuning LLMs with both theoretical guarantees and empirical evidence. Miao Lu, Shenao Zhang, Boyi Liu 0001, Hongyi Guo, Yingxiang Yang, Jose H. Blanchet, Zhaoran Wang 0001 |
NeurIPS | 2 |
| 2024 | Distributionally Robust Reinforcement Learning with Interactive Data Collection: Fundamental Hardness and Near-Optimal AlgorithmsabstractThe sim-to-real gap, which represents the disparity between training and testing environments, poses a significant challenge in reinforcement learning (RL). A promising approach to addressing this challenge is distributionally robust RL, often framed as a robust Markov decision process (RMDP). In this framework, the objective is to find a robust policy that achieves good performance under the worst-case scenario among all environments within a pre-specified uncertainty set centered around the training environment. Unlike previous work, which relies on a generative model or a pre-collected offline dataset enjoying good coverage of the deployment environment, we tackle robust RL via interactive data collection, where the learner interacts with the training environment only and refines the policy through trial and error. In this robust RL paradigm, two main challenges emerge: managing distributional robustness while striking a balance between exploration and exploitation during data collection. Initially, we establish that sample-efficient learning without additional assumptions is unattainable owing to the curse of support shift; i.e., the potential disjointedness of the distributional supports between the training and testing environments. To circumvent such a hardness result, we introduce the vanishing minimal value assumption to RMDPs with a total-variation (TV) distance robust set, postulating that the minimal value of the optimal robust value function is zero. We prove that such an assumption effectively eliminates the support shift issue for RMDPs with a TV distance robust set, and present an algorithm with a provable sample complexity guarantee. Our work makes the initial step to uncovering the inherent difficulty of robust RL via interactive data collection and sufficient conditions for designing a sample-efficient algorithm accompanied by sharp sample complexity analysis. Miao Lu, Han Zhong 0001, Tong Zhang 0001, Jose H. Blanchet |
NeurIPS | 1 |
| 2024 | Greenhouse light and CO2 regulation considering cost and photosynthesis rate using i-nsGA Ⅱ
Pan Gao 0003, Miao Lu, Yongxia Yang, Hanping Mao, Jin Hu 0007 |
Expert Syst. Appl. | 2 |
| 2023 | Pessimism in the Face of Confounders: Provably Efficient Offline Reinforcement Learning in Partially Observable Markov Decision Processes
Miao Lu, Yifei Min, Zhaoran Wang 0001, Zhuoran Yang |
ICLR | 1 |
| 2023 | Double Pessimism is Provably Efficient for Distributionally Robust Offline Reinforcement Learning: Generic Algorithm and Robust Partial CoverageabstractWe study distributionally robust offline reinforcement learning (RL), which seeks to find an optimal robust policy purely from an offline dataset that can perform well in perturbed environments. We propose a generic algorithm framework Doubly Pessimistic Model-based Policy Optimization ($\texttt{P}^2\texttt{MPO}$) for robust offline RL, which features a novel combination of a flexible model estimation subroutine and a doubly pessimistic policy optimization step. Here the double pessimism principle is crucial to overcome the distribution shift incurred by i) the mismatch between behavior policy and the family of target policies; and ii) the perturbation of the nominal model. Under certain accuracy assumptions on the model estimation subroutine, we show that $\texttt{P}^2\texttt{MPO}$ is provably sample-efficient with robust partial coverage data, which means that the offline dataset has good coverage of the distributions induced by the optimal robust policy and perturbed models around the nominal model. By tailoring specific model estimation subroutines for concrete examples including tabular Robust Markov Decision Process (RMDP), factored RMDP, and RMDP with kernel and neural function approximations, we show that $\texttt{P}^2\texttt{MPO}$ enjoys a $\tilde{\mathcal{O}}(n^{-1/2})$ convergence rate, where $n$ is the number of trajectories in the offline dataset. Notably, these models, except for the tabular case, are first identified and proven tractable by this paper. To the best of our knowledge, we first propose a general learning principle --- double pessimism --- for robust offline RL and show that it is provably efficient in the context of general function approximations. Jose H. Blanchet, Miao Lu, Tong Zhang 0001, Han Zhong 0001 |
NeurIPS | 2 |
| 2023 | Maximize to Explore: One Objective Function Fusing Estimation, Planning, and ExplorationabstractIn reinforcement learning (RL), balancing exploration and exploitation is crucial for achieving an optimal policy in a sample-efficient way. To this end, existing sample- efficient algorithms typically consist of three components: estimation, planning, and exploration. However, to cope with general function approximators, most of them involve impractical algorithmic components to incentivize exploration, such as data-dependent level-set constraints or complicated sampling procedures. To address this challenge, we propose an easy-to-implement RL framework called Maximize to Explore (MEX), which only needs to optimize unconstrainedly a single objective that integrates the estimation and planning components while balancing exploration and exploitation automatically. Theoretically, we prove that the MEX achieves a sublinear regret with general function approximators and is extendable to the zero-sum Markov game setting. Meanwhile, we adapt deep RL baselines to design practical versions of MEX in both the model-based and model-free settings, which outperform baselines in various MuJoCo environments with sparse reward by a stable margin. Compared with existing sample-efficient algorithms with general function approximators, MEX achieves similar sample efficiency while also enjoying a lower computational cost and is more compatible with modern deep RL methods. Miao Lu, Wei Xiong 0015, Han Zhong 0001, Shenao Zhang, Sirui Zheng, Zhuoran Yang, Zhaoran Wang 0001 |
NeurIPS | 2 |
| 2023 | Multi-View Multi-Task Campaign Embedding for Cold-Start Conversion Rate ForecastingabstractIn online advertising, it is critical for advertisers to forecast conversion rate (CVR) of campaigns. Previous work on campaign forecasting concentrates on the time-series analysis which depend on the availability of a length of history. However, these approaches become inadequate for cold-start campaigns which lack for the observation of past. In this work, we attempt to mitigate this challenge by learning an unsupervised and composite campaign embedding to capture multi-view semantic relationships on campaign information, and consequently forecasting the cold-start campaigns using the nearest neighbor campaigns. Specifically, we propose a novel embedding framework which simultaneously extracts and fuses heterogeneous knowledge from multiple views of campaign data in a multi-task learning fashion, to learn the semantic relationship of ad message, conversion rule, and audience targeting. We develop a hierarchical attention mechanism to refine the embedding model at two levels - an intra-view attention to improve context aggregation, and an inter-task attention to balance task importance. Finally, we adopt the k-NN regression model to predict the CVR based on the neighboring campaigns in the embedding space which encodes the multi-view campaign proximity. We conduct extensive experiments on a real-world advertising campaign dataset. The results demonstrate the effectiveness of the proposed embedding method for CVR forecasting in cold-start scenarios. Zijun Yao 0001, Deguang Kong, Miao Lu, Xiao Bai 0002, Jian Yang 0002, Hui Xiong 0001 |
IEEE Trans. Big Data | 3 |
| 2022 | Learning Robust Policy against Disturbance in Transition Dynamics via State-Conservative Policy OptimizationabstractDeep reinforcement learning algorithms can perform poorly in real-world tasks due to the discrepancy between source and target environments. This discrepancy is commonly viewed as the disturbance in transition dynamics. Many existing algorithms learn robust policies by modeling the disturbance and applying it to source environments during training, which usually requires prior knowledge about the disturbance and control of simulators. However, these algorithms can fail in scenarios where the disturbance from target environments is unknown or is intractable to model in simulators. To tackle this problem, we propose a novel model-free actor-critic algorithm---namely, state-conservative policy optimization (SCPO)---to learn robust policies without modeling the disturbance in advance. Specifically, SCPO reduces the disturbance in transition dynamics to that in state space and then approximates it by a simple gradient-based regularizer. The appealing features of SCPO include that it is simple to implement and does not require additional knowledge about the disturbance or specially designed simulators. Experiments in several robot control tasks demonstrate that SCPO learns robust policies against the disturbance in transition dynamics. Yufei Kuang, Miao Lu, Jie Wang 0005, Qi Zhou 0008, Bin Li 0025, Houqiang Li |
AAAI | 2 |
| 2022 | GEN-VLKT: Simplify Association and Enhance Interaction Understanding for HOI DetectionabstractThe task of Human-Object Interaction (HOI) detection could be divided into two core problems, i.e., human-object association and interaction understanding. In this paper, we reveal and address the disadvantages of the conventional query-driven HOI detectors from the two aspects. For the association, previous two-branch methods suffer from complex and costly post-matching, while single-branch methods ignore the features distinction in different tasks. We propose Guided-Embedding Network (GEN) to attain a two-branch pipeline without post-matching. In GEN, we design an instance decoder to detect humans and objects with two independent query sets and a position Guided Embedding (p-GE) to mark the human and object in the same position as a pair. Besides, we design an interaction decoder to classify interactions, where the interaction queries are made of instance Guided Embeddings (i-GE) generated from the outputs of each instance decoder layer. For the interaction understanding, previous methods suffer from long-tailed distribution and zero-shot discovery. This paper proposes Visual-Linguistic Knowledge Transfer (VLKT) training strategy to enhance interaction understanding by transferring knowledge from a visual-linguistic pre-trained model CLIP. In specific, we extract text embeddings for all labels with CLIP to initialize the classifier and adopt a mimic loss to minimize the visual feature distance between GEN and CLIP. As a result, GEN-VLKT outperforms the state of the art by large margins on multiple datasets, e.g., +5.05 mAP on HICO-Det. The source codes are available at https://github.com/YueLiao/gen-vlkt. Yue Liao, Aixi Zhang, Miao Lu, Si Liu 0001 |
CVPR | 3 |
| 2022 | Welfare Maximization in Competitive Equilibrium: Reinforcement Learning for Markov Exchange EconomyabstractWe study a bilevel economic system, which we refer to as a Markov exchange economy (MEE), from the point of view of multi-agent reinforcement learning (MARL). An MEE involves a central planner and a group of self-interested agents. The goal of the agents is to form a Competitive Equilibrium (CE), where each agent myopically maximizes her own utility at each step. The goal of the central planner is to steer the system so as to maximize social welfare, which is defined as the sum of the utilities of all agents. Working in a setting in which the utility function and the system dynamics are both unknown, we propose to find the socially optimal policy and the CE from data via both online and offline variants of MARL. Concretely, we first devise a novel suboptimality metric specifically tailored to MEE, such that minimizing such a metric certifies globally optimal policies for both the planner and the agents. Second, in the online setting, we propose an algorithm, dubbed as \texttt{MOLM}, which combines the optimism principle for exploration with subgame CE seeking. Our algorithm can readily incorporate general function approximation tools for handling large state spaces and achieves a sublinear regret. Finally, we adapt the algorithm to an offline setting based on the pessimism principle and establish an upper bound on the suboptimality. Miao Lu, Zhaoran Wang 0001, Michael I. Jordan, Zhuoran Yang |
ICML | 2 |
| 2021 | Mining the Benefits of Two-stage and One-stage HOI DetectionabstractTwo-stage methods have dominated Human-Object Interaction~(HOI) detection for several years. Recently, one-stage HOI detection methods have become popular. In this paper, we aim to explore the essential pros and cons of two-stage and one-stage methods. With this as the goal, we find that conventional two-stage methods mainly suffer from positioning positive interactive human-object pairs, while one-stage methods are challenging to make an appropriate trade-off on multi-task learning, \emph{i.e.}, object detection, and interaction classification. Therefore, a core problem is how to take the essence and discard the dregs from the conventional two types of methods. To this end, we propose a novel one-stage framework with disentangling human-object detection and interaction classification in a cascade manner. In detail, we first design a human-object pair generator based on a state-of-the-art one-stage HOI detector by removing the interaction classification module or head and then design a relatively isolated interaction classifier to classify each human-object pair. Two cascade decoders in our proposed framework can focus on one specific task, detection or interaction classification. In terms of the specific implementation, we adopt a transformer-based HOI detector as our base model. The newly introduced disentangling paradigm outperforms existing methods by a large margin, with a significant relative mAP gain of 9.32% on HICO-Det. The source codes are available at https://github.com/YueLiao/CDN. Aixi Zhang, Yue Liao, Si Liu 0001, Miao Lu, Chen Gao 0005 |
NeurIPS | 4 |
| 2020 | First Impression: AI Understands PersonalityabstractWhen you first encounter a person, a mental image of that person is formed. First impression, an interactive art, is proposed to let AI understand human personality at first glance. The mental image is demonstrated by Beijing opera facial makeups, which shows the character personality with a combination of realism and symbolism. We build Beijing opera facial makeup dataset and semantic dataset of facial features to establish relationships among real faces, personalities and facial makeups. First impression detects faces, recognizes personality from facial appearance and finds the matching Beijing opera facial makeup. Finally, the morphing process from real face to facial makeup is shown to let users enjoy the process of AI understanding personality. Xiaohui Wang 0004, Xia Liang, Miao Lu, Jingyan Qin |
ACM Multimedia | 3 |
| 2020 | Learning from Cross-Modal Behavior Dynamics with Graph-Regularized Neural Contextual BanditabstractContextual multi-armed bandit algorithms have received significant attention in modeling users’ preferences for online personalized recommender systems in a timely manner. While significant progress has been made along this direction, a few major challenges have not been well addressed yet: (i) a vast majority of the literature is based on linear models that cannot capture complex non-linear inter-dependencies of user-item interactions; (ii) existing literature mainly ignores the latent relations among users and non-recommended items: hence may not properly reflect users’ preferences in the real-world; (iii) current solutions are mainly based on historical data and are prone to cold-start problems for new users who have no interaction history. Xian Wu 0003, Suleyman Cetintas, Deguang Kong, Miao Lu, Jian Yang 0002, Nitesh V. Chawla |
WWW | 4 |
| 2019 | A Quaternion's Encoding Sine Cosine Algorithm
Dengxu He, Miao Lu, Yundi Rao |
ICIC (1) | 3 |
| 2018 | SACP: A Signcryption-Based Authentication Scheme with Conditional Privacy Preservation for VANET
Miao Lu, Ying Wu 0006 |
WASA | 1 |
| 2018 | Attention Convolutional Neural Network for Advertiser-level Click-through Rate ForecastingabstractClick-through rate (CTR) is a critical problem in online advertising. Most existing researches only focus on the user-level CTR prediction. However, advertiser-level CTR forecasting also plays a very important role because advertisers typically decide how much they would like to bid for advertisements to achieve the maximum clicks given their budget based on CTR forecasting. Over-forecasting will make the advertiser to pay more than necessary but get less return on investment (ROI). Under-forecasting will make the advertiser to spend less money on campaigns but they cannot achieve the desired ROI goals. In this paper, we focus on the advertiser-level CTR forecasting and formulate it as a time series forecasting problem based on the historical CTR record. This is a very challenging problem due to the heavy fluctuation and highly non-linearity of time series. Furthermore, advertisers usually provide useful contextual information for their campaigns, such as text descriptions, targeting locations and devices, which has high correlation with CTR but has not yet been used for CTR forecasting. Thus, we propose a novel context-aware attention convolutional neural network (CACNN), which can capture the high non-linearity and local information of the time series, as well as the underlying correlation between the time series of CTR and the contextual information. To the best of our knowledge, this is the first work employing convolutional neural network and incorporating heterogeneous information to perform CTR forecasting at advertiser level. We implement the system on Yahoo TensorFlowOnSpark platform which enables distributed deep learning on a cluster of GPU and CPU servers, and achieves faster learning speed and data access on HDFS when available. The effectiveness of CACNN model has been demonstrated in real-world Yahoo advertising dataset, and therefore deployed in production with daily rolling of the model. Hongchang Gao, Deguang Kong, Miao Lu, Xiao Bai 0002, Jian Yang 0002 |
WWW | 3 |
| 2016 | A Calibration-Free Crowdsourcing-Based Indoor Localization Solution
Ying Wu 0006, Miao Lu |
APSCC | 4 |
| 2016 | Extending the Pairwise Separability Index for Multicrop Identification Using Time-Series MODIS ImagesabstractThe pairwise separability index (SI) has been demonstrated as an effective indicator for capturing crucial phenological differences between two plant species. However, its application to crop types, which have more obvious phenological characteristics than natural vegetation, has received less attention, and extending the pairwise SI to multiple crops for feature selection still remains a challenge. This paper presented two SI extension approaches (SIaveand SImin) to select the optimal spectro-temporal features for multiple crops, and investigated their classification performance using Heilongjiang Province, China, as a study area. Feature interpretability and classification accuracy of different crops were evaluated for the two approaches. The results showed that the SIaveapproach generally has relatively high feature interpretability due to its better description of crucial phenological characteristics of different crops. Those crops with high separability are insensitive to the extension approach and have similar classification accuracy for the two approaches, whereas those crops with poor separability show good performance with the SIminmethod. Due to the higher temporal autocorrelation, the optimal features for crop classification that are selected by the SIaveapproach exhibit greater information redundancy across the time domain than those that are selected by the SIminapproach, which largely explains the relatively low classification accuracy achieved using the SIaveapproach. These comparison results between SIminand SIaveapproaches also indicate that time-series images with high temporal resolution do not necessarily produce high classification accuracy, regardless of their ability to describe the seasonal characteristics of crops. Qiong Hu 0002, Qian Song, Qiangyi Yu, Miao Lu, Peng Yang 0005, Huajun Tang, Yuqiao Long |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2015 | Automatic change detection of urban land-cover based on SVM classificationabstractThe reliability of support vector machines for classifying multi-spectral images of remote sensing has been proven in various studies. In this paper, we investigate their applicability for urban land cover in Wuhan, Hubei province of China. Firstly, radiation rectification, normalization processing and geometry registration are made between the bi-temporal images. Secondly, SVM approach is used in our study to classify sorts and land use types from bi-temporal images. Thirdly, build matrix of change detection in basis of the potential types of change. Post-classification compare are proposed pixel-by-pixel. According to the sort of change of every pixel, new value is assigned on the base of change matrix. The output is image of change. Lastly, the process and pattern of the urban land use change in the Wuhan district was finally revealed from 2009 to 2013 in our study. Miao Lu, Xiuwan Chen |
IGARSS | 2 |
| 2008 | A Data-Aided Channel Estimation Method Based on CAZACabstractFrequency-hopped OFDM (FH-OFDM) technology will provide the additional benefits of frequency diversity and inter-cell interference averaging. However, it is difficult to get the channel state information. In this paper, we propose a method of channel estimation over slow fading channels for FH-OFDM systems. In the algorithm, the channel impulse response is estimated by correlation operation, by using the cyclic shift zero- correlation of CAZAC (constant amplitude zero auto-correlation) sequence. In FH-OFDM systems, the influence of noise on channel estimation precision is restrained and inter-cell users' interference can be eliminated. At the same time, the algorithm is adapted to any hopping pattern. Miao Lu, Xiaolin Hou, Dongmei Luo |
VTC Fall | 1 |
| 2007 | Analysis and Simulation for Radio Access Network Architecture of 3GPP Long Term EvolutionabstractIn this paper, two popular radio access network architectures, centralized architecture and distributed architecture, are analyzed. Then we propose a novel architecture on the basis of the comparisons and system level simulation results. In this new architecture, control plane and user plane are terminated in different nodes and can be processed simultaneously. Analysis and simulation results show that this proposed network architecture can provide better performance in lower signaling overhead and robustness. Haipeng Lei, Yafeng Wang, Miao Lu, Dacheng Yang |
PIMRC | 3 |
| 2007 | Study on EUTRA DFT-S OFDM Uplink Channel EstimationabstractLocalized transmission with frequency hopping in E-UTRA uplink can avoid some drawbacks of distributed transmission and offer improved possibilities for the frequency diversity. However, it is difficult to find out a general hopping pattern applicable to various transmission bandwidths while keeping single carrier transmission at the same time. In this paper, a new RB (Resource Block) adaptive hopping method based on predefined pattern of intra/inter-TTI frequency hopping is proposed for E-UTRA uplink. The method that can keep single carrier transmission and avoid collisions is suitable for various transmission bandwidths. Simulation results show that the proposed hopping mode is effective. Miao Lu, Yafeng Wang, Haipeng Lei, Dacheng Yang |
PIMRC | 1 |
| 2006 | A Differential Evolution Based Method for Power System PlanningabstractPower system planning is a complex multi-objective optimization problem. It aims at locating the minimum cost of additional transmission lines that must be installed to satisfy the forecasted load in a power system. A number of different methods for power system planning have been investigated over the past decades. In this paper, a differential evolution (DE) based approach is proposed as an optimization tool to solve the power system planning problem. A comparison between genetic algorithms, evolutionary strategy (ES), and five different DE schemes are carried out on two benchmark power systems. The results shown that, as a relatively new heuristic optimization method, DE is able to provide robust and efficient solution to power system planning problems. Zhao Yang Dong, Miao Lu, Zhe Lu, Kit Po Wong |
IEEE Congress on Evolutionary Computation | 2 |