EDBT 2026 Demo / reviewers in the wild / expert
Qingyang Li 0001
dblp:70/11398-1
· DBLP profile ↗
22ranked-venue papers
3as first author
14since 2021 · last 2026
0000-0001-6650-2343ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 12 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | OxygenREC: An Instruction-Following Generative Framework for E-commerce Recommendation
Qingyang Li 0001, Yanchen Qiao, Ziyang Ji, Xiangyu Qian, Yanlong Zang, Weijie Ding, Yaqiang Zang, Pinghua Gong |
SIGIR | 2 |
| 2025 | Towards Reward Fairness in RLHF: From a Resource Allocation PerspectiveabstractRewards serve as proxies for human preferences and play a crucial role in Reinforcement Learning from Human Feedback (RLHF).However, if these rewards are inherently imperfect, exhibiting various biases, they can adversely affect the alignment of large language models (LLMs).In this paper, we collectively define the various biases present in rewards as the problem of reward unfairness.We propose a bias-agnostic method to address the issue of reward fairness from a resource allocation perspective, without specifically designing for each type of bias, yet effectively mitigating them.Specifically, we model preference learning as a resource allocation problem, treating rewards as resources to be allocated while considering the trade-off between utility and fairness in their distribution.We propose two methods, Fairness Regularization and Fairness Coefficient, to achieve fairness in rewards.We apply our methods in both verification and reinforcement learning scenarios to obtain a fairness reward model and a policy model, respectively.Experiments conducted in these scenarios demonstrate that our approach aligns LLMs with human preferences in a more fair manner.Our data and code are available at https://github.com/ shoyua/Towards-Reward-Fairness. Sheng Ouyang, Yulan Hu, Ge Chen 0006, Qingyang Li 0001, Yong Liu 0018 |
ACL (1) | 4 |
| 2024 | Fewer is more: efficient object detection in large aerial images
Xingxing Xie, Gong Cheng 0003, Qingyang Li 0001, Shicheng Miao, Ke Li 0005, Junwei Han 0001 |
Sci. China Inf. Sci. | 3 |
| 2023 | Sim2Rec: A Simulator-based Decision-making Approach to Optimize Real-World Long-term User Engagement in Sequential Recommender SystemsabstractLong-term user engagement (LTE) optimization in sequential recommender systems (SRS) is shown to be suited by reinforcement learning (RL) which finds a policy to maximize long-term rewards. Meanwhile, RL has its shortcomings, particularly requiring a large number of online samples for exploration, which is risky in real-world applications. One of the appealing ways to avoid the risk is to build a simulator and learn the optimal recommendation policy in the simulator. In LTE optimization, the simulator is to simulate multiple users’ daily feedback for given recommendations. However, building a user simulator with no reality-gap, i.e., can predict user’s feedback exactly, is unrealistic because the users’ reaction patterns are complex and historical logs for each user are limited, which might mislead the simulator-based recommendation policy. In this paper, we present a practical simulator-based recommender policy training approach, Simulation-to-Recommendation (Sim2Rec) to handle the reality-gap problem for LTE optimization. Specifically, Sim2Rec introduces a simulator set to generate various possibilities of user behavior patterns, then trains an environment-parameter extractor to recognize users’ behavior patterns in the simulators. Finally, a context-aware policy is trained to make the optimal decisions on all of the variants of the users based on the inferred environment-parameters. The policy is transferable to unseen environments (e.g., the real world) directly as it has learned to recognize all various user behavior patterns and to make the correct decisions based on the inferred environment-parameters. Experiments are conducted in synthetic environments and a real-world large-scale ride-hailing platform, DidiChuxing. The results show that Sim2Rec achieves significant performance improvement, and produces robust recommendations in unseen environments. Xiong-Hui Chen, Bowei He, Yang Yu 0001, Qingyang Li 0001, Zhiwei (Tony) Qin, Wenjie Shang, Jieping Ye, Chen Ma 0001 |
ICDE | 4 |
| 2023 | Offline Model-Based Adaptable Policy Learning for Decision-Making in Out-of-Support RegionsabstractIn reinforcement learning, a promising direction to avoid online trial-and-error costs is learning from an offline dataset. Current offline reinforcement learning methods commonly learn in the policy space constrained to in-support regions by the offline dataset, in order to ensure the robustness of the outcome policies. Such constraints, however, also limit the potential of the outcome policies. In this paper, to release the potential of offline policy learning, we investigate the decision-making problems in out-of-support regions directly and propose offline Model-based Adaptable Policy LEarning (MAPLE). By this approach, instead of learning in in-support regions, we learn an adaptable policy that can adapt its behavior in out-of-support regions when deployed. We give a practical implementation of MAPLE via meta-learning techniques and ensemble model learning techniques. We conduct experiments on MuJoCo locomotion tasks with offline datasets. The results show that the proposed method can make robust decisions in out-of-support regions and achieve better performance than SOTA algorithms. Xiong-Hui Chen, Fan-Ming Luo, Yang Yu 0001, Qingyang Li 0001, Zhiwei (Tony) Qin, Wenjie Shang, Jieping Ye |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | SFRNet: Fine-Grained Oriented Object Recognition via Separate Feature RefinementabstractFine-grained oriented object recognition (FGO2R) is a practical need for intellectually interpreting remote sensing images. It aims at realizing fine-grained classification and precise localization with oriented bounding boxes, simultaneously. Our considerations for the task are general but decisive: (i) the extraction of subtle differences carries a big weight in differentiating fine-grained classes, and (ii) oriented localization prefers rotation-sensitive features. In this article, we propose a network with separate feature refinement (SFRNet), in which two transformer-based branches are designed to perform function-specific feature refinement for fine-grained classification and oriented localization, separately. To highlight the discriminative information advantageous to fine-grained classification, we propose a spatial and channel transformer (SC-Former) to capture both the long-range spatial interactions and the key correlations hidden in the feature channels. Besides, we design a Multi-RoI loss (MRL) following the protocol of deep metric learning to enhance the separability of fine-grained classes further. For oriented localization, we integrate the oriented response convolution with the transformer structure (namely, OR-Former) to assist in encoding rotation information during regression. Extensive experimental results validate the effectiveness and robustness of our SFRNet. Without bells and whistles, our SFRNet achieves state-of-the-art performance on the large-scale FAIR1M datasets (FAIR1M-1.0 and FAIR1M-2.0). Code will be available at https://github.com/Ranchosky/SFRNet. Gong Cheng 0003, Qingyang Li 0001, Guangxing Wang 0001, Xingxing Xie, Lingtong Min, Junwei Han 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Multiple Tiered Treatments Optimization with Causal Inference on Response DistributionabstractFor many business applications, event driven promotion programs are commonly used means for achieving business targets and performances such as customer experience and revenue etc. Traditional recommendation methods mainly focus on predicting the click-through-rate and can not handle this scenario well since most of driven promotion are multi-tiered programs. Moreover, promotion programs will incur costs, so we need to predict not only the effect of recommending programs to users, but also the cost of programs. So it is essentially an NP-hard optimization problem with a budget constraint. Recently, there are some studies that conducts causal inferences for analyzing the effects of various designed programs on users. However, causal inference models mostly address the expected effects of treatments instead of studying the probabilistic distribution of the heterogeneous effects on users, thus unable to help on accurate estimation of the program costs driven by events. In this paper, we argue that the expected treatment effects can be learned with the response distribution, and proposed a multi-task learning model for composite causal inference on both treatment effects and response distribution, which allows decision makers to make trade-off decisions on the return and cost of programs in an end-to-end manner and to optimize their policy and strategy. Qingyang Li 0001, Zhiwei (Tony) Qin |
IEEE Big Data | 2 |
| 2022 | Dynamic Proposal Generation for Oriented Object Detection in Aerial ImagesabstractCurrent two-stage oriented object detectors for aerial images have achieved remarkable progress. However, they still suffer from some drawbacks. Firstly, most of them place redundant anchors or utilize complicated transformation to generate oriented proposals, which are inefficient. Secondly, the generation of proposals is static, which cannot adapt to the extremely nonuniform distribution of objects. To address these issues, we propose a Dynamic Proposal Generation Network (DPGN) which can generate high-quality oriented proposals directly and estimate the upper limit of proposals adaptively. To be specific, with Guided Anchor Regression (GAR), we obtain the coarse oriented anchors and utilize them to align the features. After this, we make further classification and regression to produce final oriented proposals. Meanwhile, we design Maximum Number Estimation (MNE) for predicting an approximate value to remain the proposals adaptively. Without tricks, our method can achieve competitive detection accuracy compared with other mainstream methods on DOTA dataset. Qingyang Li 0001, Gong Cheng 0003, Shicheng Miao |
IGARSS | 1 |
| 2022 | Precise Vertex Regression and Feature Decoupling for Oriented Object DetectionabstractOriented object detection is a key task in the field of remote sensing image interpretation. Although extensive efforts have been made over the past few years, accurate oriented object detection remains a big challenge due to the dense arrangement and diverse orientations of objects. In this paper, we propose an oriented object detector based on the Faster R-CNN, which mainly consists of a Precise Vertex Regression (PVR) module and a Feature Decoupling (FD) module. Specifically, the PVR module predicts the arbitrary quadrilaterals of oriented objects with the precise vertex regression manner, which discretizes the regression range of vertex into several bins and applies a classification network to predict which bin the vertex belongs to. The FD module decouples the RoI features for classification and regression tasks by lightweight affine transformation. Experimental results on DOTA and DIOR-R datasets validate the effectiveness of our proposed method. Code is available at https://github.com/ShichengMiao16/VRDet. Shicheng Miao, Gong Cheng 0003, Qingyang Li 0001 |
IGARSS | 3 |
| 2021 | Multi-Scale Bidirectional Feature Fusion for One-Stage Oriented Object Detection in Aerial ImagesabstractThis paper aims to address the problem of oriented object detection under the complex background of remote sensing images. To this end, we propose a one-stage object detection method with feature fusion structure, and modify the loss function to enhance the detection of small objects. More specifically, on the basis of the end-to-end one-stage object detection model RetinaNet, the method of gliding the vertices of the horizontal bounding box is used to describe an oriented object. In order to obtain multi-scale context information, we design a feature fusion module. Besides, we propose a novel area-weighted loss function to pay more attention to small objects. Experimental results conducted on the DOTA dataset demonstrate that the proposed framework outperforms several state-of-the-art baselines. Gong Cheng 0003, Xuxiang Sun 0001, Qingyang Li 0001, Meili Zhang, Shicheng Miao |
IGARSS | 4 |
| 2021 | Offline Model-based Adaptable Policy LearningabstractIn reinforcement learning, a promising direction to avoid online trial-and-error costs is learning from an offline dataset. Current offline reinforcement learning methods commonly learn in the policy space constrained to in-support regions by the offline dataset, in order to ensure the robustness of the outcome policies. Such constraints, however, also limit the potential of the outcome policies. In this paper, to release the potential of offline policy learning, we investigate the decision-making problems in out-of-support regions directly and propose offline Model-based Adaptable Policy LEarning (MAPLE). By this approach, instead of learning in in-support regions, we learn an adaptable policy that can adapt its behavior in out-of-support regions when deployed. We conduct experiments on MuJoCo controlling tasks with offline datasets. The results show that the proposed method can make robust decisions in out-of-support regions and achieve better performance than SOTA algorithms. Xiong-Hui Chen, Yang Yu 0001, Qingyang Li 0001, Fan-Ming Luo, Zhiwei (Tony) Qin, Wenjie Shang, Jieping Ye |
NeurIPS | 3 |
| 2021 | Predicting future cognitive decline with hyperbolic stochastic coding
Jie Zhang 0026, Qunxi Dong, Jie Shi 0001, Qingyang Li 0001, Cynthia M. Stonnington, Boris Gutman, Kewei Chen 0001, Eric Reiman, Richard J. Caselli, Paul M. Thompson, Jieping Ye, Yalin Wang 0001 |
Medical Image Anal. | 4 |
| 2021 | Partially observable environment estimation with uplift inference for reinforcement learning based recommendationabstractReinforcement learning (RL) aims at searching the best policy model for decision making, and has been shown powerful for sequential recommendations. The training of the policy by RL, however, is placed in an environment. In many real-world applications, the policy training in the real environment can cause an unbearable cost due to the exploration. Environment estimation from the past data is thus an appealing way to release the power of RL in these applications. The estimation of the environment is, basically, to extract the causal effect model from the data. However, real-world applications are often too complex to offer fully observable environment information. Therefore, quite possibly there are unobserved variables lying behind the data, which can obstruct an effective estimation of the environment. In this paper, by treating the hidden variables as a hidden policy, we propose a partially-observed multi-agent environment estimation (POMEE) approach to learn the partially-observed environment. To make a better extraction of the causal relationship between actions and rewards, we design a deep uplift inference network (DUIN) model to learn the causal effects of different actions. By implementing the environment model in the DUIN structure, we propose a POMEE with uplift inference (POMEE-UI) approach to generate a partially-observed environment with a causal reward mechanism. We analyze the effect of our method in both artificial and real-world environments. We first use an artificial recommender environment, abstracted from a real-world application, to verify the effectiveness of POMEE-UI. We then test POMEE-UI in the real application of Didi Chuxing. Experiment results show that POMEE-UI can effectively estimate the hidden variables, leading to a more reliable virtual environment. The online A/B testing results show that POMEE can derive a well-performing recommender policy in the real-world application. Wenjie Shang, Qingyang Li 0001, Zhiwei (Tony) Qin, Yang Yu 0001, Yiping Meng, Jieping Ye |
Mach. Learn. | 2 |
| 2021 | Multi-Resemblance Multi-Target Low-Rank Coding for Prediction of Cognitive Decline With Longitudinal Brain ImagesabstractAn effective presymptomatic diagnosis and treatment of Alzheimer's disease (AD) would have enormous public health benefits. Sparse coding (SC) has shown strong potential for longitudinal brain image analysis in preclinical AD research. However, the traditional SC computation is time-consuming and does not explore the feature correlations that are consistent over the time. In addition, longitudinal brain image cohorts usually contain incomplete image data and clinical labels. To address these challenges, we propose a novel two-stage Multi-Resemblance Multi-Target Low-Rank Coding (MMLC) method, which encourages that sparse codes of neighboring longitudinal time points are resemblant to each other, favors sparse code low-rankness to reduce the computational cost and is resilient to both source and target data incompleteness. In stage one, we propose an online multi-resemblant low-rank SC method to utilize the common and task-specific dictionaries in different time points to immune to incomplete source data and capture the longitudinal correlation. In stage two, supported by a rigorous theoretical analysis, we develop a multi-target learning method to address the missing clinical label issue. To solve such a multi-task low-rank sparse optimization problem, we propose multi-task stochastic coordinate coding with a sequence of closed-form update steps which reduces the computational costs guaranteed by a theoretical convergence proof. We apply MMLC on a publicly available neuroimaging cohort to predict two clinical measures and compare it with six other methods. Our experimental results show our proposed method achieves superior results on both computational efficiency and predictive accuracy and has great potential to assist the AD prevention. Jie Zhang 0026, Qingyang Li 0001, Richard J. Caselli, Paul M. Thompson, Jieping Ye, Yalin Wang 0001 |
IEEE Trans. Medical Imaging | 3 |
| 2020 | Hierarchical Adaptive Contextual Bandits for Resource Constraint based RecommendationabstractContextual multi-armed bandit (MAB) achieves cutting-edge performance on a variety of problems. When it comes to real-world scenarios such as recommendation system and online advertising, however, it is essential to consider the resource consumption of exploration. In practice, there is typically non-zero cost associated with executing a recommendation (arm) in the environment, and hence, the policy should be learned with a fixed exploration cost constraint. It is challenging to learn a global optimal policy directly, since it is a NP-hard problem and significantly complicates the exploration and exploitation trade-off of bandit algorithms. Existing approaches focus on solving the problems by adopting the greedy policy which estimates the expected rewards and costs and uses a greedy selection based on each arm’s expected reward/cost ratio using historical observation until the exploration resource is exhausted. However, existing methods are hard to extend to infinite time horizon, since the learning process will be terminated when there is no more resource. In this paper, we propose a hierarchical adaptive contextual bandit method (HATCH) to conduct the policy learning of contextual bandits with a budget constraint. HATCH adopts an adaptive method to allocate the exploration resource based on the remaining resource/time and the estimation of reward distribution among different user contexts. In addition, we utilize full of contextual feature information to find the best personalized recommendation. Finally, in order to prove the theoretical guarantee, we present a regret bound analysis and prove that HATCH achieves a regret bound as low as . The experimental results demonstrate the effectiveness and efficiency of the proposed method on both synthetic data sets and the real-world applications. Mengyue Yang, Qingyang Li 0001, Zhiwei (Tony) Qin, Jieping Ye |
WWW | 2 |
| 2019 | Environment Reconstruction with Hidden Confounders for Reinforcement Learning based RecommendationabstractReinforcement learning aims at searching the best policy model for decision making, and has been shown powerful for sequential recommendations. The training of the policy by reinforcement learning, however, is placed in an environment. In many real-world applications, however, the policy training in the real environment can cause an unbearable cost, due to the exploration in the environment. Environment reconstruction from the past data is thus an appealing way to release the power of reinforcement learning in these applications. The reconstruction of the environment is, basically, to extract the casual effect model from the data. However, real-world applications are often too complex to offer fully observable environment information. Therefore, quite possibly there are unobserved confounding variables lying behind the data. The hidden confounder can obstruct an effective reconstruction of the environment. In this paper, by treating the hidden confounder as a hidden policy, we propose a deconfounded multi-agent environment reconstruction (DEMER) approach in order to learn the environment together with the hidden confounder. DEMER adopts a multi-agent generative adversarial imitation learning framework. It proposes to introduce the confounder embedded policy, and use the compatible discriminator for training the policies. We then apply DEMER in an application of driver program recommendation. We firstly use an artificial driver program recommendation environment, abstracted from the real application, to verify and analyze the effectiveness of DEMER. We then test DEMER in the real application of Didi Chuxing. Experiment results show that DEMER can effectively reconstruct the hidden confounder, and thus can build the environment better. DEMER also derives a recommendation policy with a significantly improved performance in the test phase of the real application. Wenjie Shang, Yang Yu 0001, Qingyang Li 0001, Zhiwei (Tony) Qin, Yiping Meng, Jieping Ye |
KDD | 3 |
| 2017 | Task fMRI data analysis based on supervised stochastic coordinate coding
Jinglei Lv, Qingyang Li 0001, Wei Zhang 0090, Yu Zhao 0007, Xi Jiang 0001, Lei Guo 0002, Junwei Han 0001, Xintao Hu, Christine Cong Guo, Jieping Ye, Tianming Liu 0001 |
Medical Image Anal. | 3 |
| 2016 | Parallel Lasso Screening for Big Data OptimizationabstractLasso regression is a widely used technique in data mining for model selection and feature extraction. In many applications, it remains challenging to apply the regression model to large-scale problems that have massive data samples with high-dimensional features. One popular and promising strategy is to solve the Lasso problem in parallel. Parallel solvers run multiple cores in parallel on a shared memory system to speedup the computation, while the practical usage is limited by the huge dimension in the feature space. Screening is a promising method to solve the problem of high dimensionality by discarding the inactive features and removing them from optimization. However, when integrating screening methods with parallel solvers, most of solvers cannot guarantee the convergence on the reduced feature matrix. In this paper, we propose a novel parallel framework by parallelizing screening methods and integrating it with our proposed parallel solver. We propose two parallel screening algorithms: Parallel Strong Rule (PSR) and Parallel Dual Polytope Projection (PDPP). For the parallel solver, we proposed an Asynchronous Grouped Coordinate Descent method (AGCD) to optimize the regression problem in parallel on the reduced feature matrix. AGCD is based on a grouped selection strategy to select the coordinate that has the maximum descent for the objective function in a group of candidates. Empirical studies on the real-world datasets demonstrate that the proposed parallel framework has a superior performance compared to the state-of-the-art parallel solvers. Qingyang Li 0001, Shuiwang Ji, Paul M. Thompson, Jieping Ye, Jie Wang 0005 |
KDD | 1 |
| 2016 | Hyperbolic Space Sparse Coding with Its Application on Prediction of Alzheimer's Disease in Mild Cognitive Impairment
Jie Zhang 0026, Jie Shi 0001, Cynthia M. Stonnington, Qingyang Li 0001, Boris Gutman, Kewei Chen 0001, Eric Reiman, Richard J. Caselli, Paul M. Thompson, Jieping Ye, Yalin Wang 0001 |
MICCAI (1) | 4 |
| 2016 | Large-Scale Collaborative Imaging Genetics Studies of Risk Genetic Factors for Alzheimer's Disease Across Multiple Institutions
Qingyang Li 0001, Tao Yang 0016, Liang Zhan, Derrek P. Hibar, Neda Jahanshad, Yalin Wang 0001, Jieping Ye, Paul M. Thompson, Jie Wang 0005 |
MICCAI (1) | 1 |
| 2014 | A Highly Scalable Parallel Algorithm for Isotropic Total Variation ModelsabstractTotal variation (TV) models are among the most popular and successful tools in signal processing. However, due to the complex nature of the TV term, it is challenging to efficiently compute a solution for large-scale problems. State-of-the-art algorithms that are based on the alternating direction method of multipliers (ADMM) often involve solving large-size linear systems. In this paper, we propose a highly scalable parallel algorithm for TV models that is based on a novel decomposition strategy of the problem domain. As a result, the TV models can be decoupled into a set of small and independent subproblems, which admit closed form solutions. This makes our approach particularly suitable for parallel implementation. Our algorithm is guaranteed to converge to its global minimum. With N variables and n_p processes, the time complexity is O(N/(εn_p)) to reach an epsilon-optimal solution. Extensive experiments demonstrate that our approach outperforms existing state-of-the-art algorithms, especially in dealing with high-resolution, mega-size images. Jie Wang 0005, Qingyang Li 0001, Sen Yang 0004, Wei Fan 0001, Peter Wonka, Jieping Ye |
ICML | 2 |
| 2013 | Adaptive Fault Detection for Testing Tenant Applications in Multi-tenancy SaaS SystemsabstractSaaS (Software-as-a-Service) often uses multi-tenancy architecture (MTA) where tenant developers compose their applications online using the components stored in the SaaS database. Tenant applications need to be tested, and combinatorial testing can be used. While numerous combinatorial testing techniques are available, most of them produce static sequences of test configurations and their goal is often to provide sufficient coverage such as 2-way interaction coverage. But the goal of SaaS testing is to identify those compositions that are faulty for tenant applications. This paper proposes an adaptive test configuration generation algorithm AR (Adaptive Reasoning) that can rapidly identify those faulty combinations so that those faulty combinations cannot be selected by tenant developers for composition. The AR algorithm has been evaluated by both simulation and real experimentation using a MTA SaaS sample running on GAE (Google App Engine). Both the simulation and experiment showed show that the AR algorithm can identify those faulty combinations rapidly. Whenever a new component is submitted to the SaaS database, the AR algorithm can be applied so that any faulty interactions with new components can be identified to continue to support future tenant applications. Wei-Tek Tsai, Qingyang Li 0001, Charles J. Colbourn, Xiaoying Bai |
IC2E | 2 |