EDBT 2026 Demo / reviewers in the wild / expert
Tianqi Shi
dblp:240/8461
· DBLP profile ↗
16ranked-venue papers
3as first author
15since 2021 · last 2026
0000-0003-4815-4175ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021Security and privacy · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DetectRL-X: Towards Reliable Multilingual and Real-World LLM-Generated Text DetectionabstractThe effective detection and governance of Large Language Model (LLM) generated content has become increasingly critical due to the growing risk of misuse. Despite the impressive performance of existing detectors, their reliability and potential in multilingual, real-world scenarios remain largely underexplored.In this study, we introduce DetectRL-X, a comprehensive multilingual benchmark designed to evaluate advanced detectors across 8 dimensions. The benchmark encompasses 8 languages commonly used in commercial contexts and collects human-written texts from 6 domains highly susceptible to LLM misuse. To better aligned with real-world applications, We create LLM-generated texts using 4 popular commercial LLMs, and include typical AI-assisted writing operations such as polishing, expanding, and condensing to capture authentic usage patterns. Furthermore, we develop a multilingual framework for paraphrasing and perturbation attacks to simulate diverse human modifications and writing noise, enabling stress testing of detectors across languages.Experimental results on DetectRL-X reveal the strengths and limitations of current state-of-the-art detectors when applied to diverse linguistic resources. We further analyze how domains, generators, attack strategies, text length, and refinement operations influence performance in different languages, underscoring DetectRL-X as an effective benchmark for strengthening multilingual and language-specific detectors. Junchao Wu, Yefeng Liu, Chenyu Zhu, Tianqi Shi, Yichao Du, Longyue Wang, Weihua Luo, Jinsong Su, Derek F. Wong |
ACL (1) | 6 |
| 2025 | Marco-o1 v2: Towards Widening The Distillation Bottleneck for Reasoning ModelsabstractLarge Reasoning Models (LRMs) such as OpenAI o1 and DeepSeek-R1 have shown remarkable reasoning capabilities by scaling test-time compute and generating long Chain-of-Thought (CoT). Distillation post-training on LRMs-generated data is a straightforward yet effective method to enhance the reasoning abilities of smaller models, but faces a critical bottleneck: we found that distilled long CoT data poses learning difficulty for small models and leads to the inheritance of biases (i.e., formalistic long-time thinking) when using Supervised Fine-tuning (SFT) and Reinforcement Learning (RL) methods. To alleviate this bottleneck, we propose constructing data from scratch using Monte Carlo Tree Search (MCTS). We then exploit a set of CoT-aware approaches, including Thoughts Length Balance, Fine-grained DPO, and Joint Post-training Objective, to enhance SFT and RL on the MCTS data. We conducted evaluation on various benchmarks such as math (GSM8K, MATH, AIME). instruction-following (Multi-IF) and planning (Blocksworld), results demonstrate our CoT-aware approaches substantially improve the reasoning performance of distilled models compared to standard distilled models via reducing the hallucinations in long-time thinking. Huifeng Yin, Minghao Wu, Xuanfan Ni, Tianqi Shi, Liangying Shao, Chenyang Lyu, Longyue Wang, Weihua Luo, Kaifu Zhang |
ACL (1) | 7 |
| 2025 | Marco-Bench-MIF: On Multilingual Instruction-Following Capability of Large LanguageabstractBo Zeng, Chenyang Lyu, Sinuo Liu, Mingyan Zeng, Minghao Wu, Xuanfan Ni, Tianqi Shi, Yu Zhao, Yefeng Liu, Chenyu Zhu, Ruizhe Li, Jiahui Geng, Qing Li, Yu Tong, Longyue Wang, Weihua Luo, Kaifu Zhang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Chenyang Lyu, Sinuo Liu, Mingyan Zeng, Minghao Wu, Xuanfan Ni, Tianqi Shi, Yefeng Liu, Chenyu Zhu, Ruizhe Li 0001, Jiahui Geng, Longyue Wang, Weihua Luo, Kaifu Zhang |
ACL (1) | 7 |
| 2025 | KHAD: K-Hop and Activation-aware Defense against Bit-Flip Attacks in Deep Neural NetworksabstractFor quantized deep neural networks widely deployed on hardware-accelerated platforms, bit-flip attacks (BFAs) have become a serious security threat because they can cripple or hijack model inference by modifying only a few bits. To this end, we propose KHAD (K-Hop and Activation-aware Defense), a unified, minimally intrusive defense framework that fuses k-hop propagation with activation statistics to precisely assess neuronal criticality and adaptively partitions neurons into three categories, to which it applies lightweight strategies: elastic boundary rectification, soft limiting with least significant bit masking, and range tightening with orthogonal diffusion, respectively. We conduct performance experiments across datasets of different scales and multiple models. The results show that, while exerting only a very small impact on clean accuracy (e.g., a decrease of 1.73% for ResNet-32 on CIFAR-100), KHAD exhibits strong defensive capability. Under untargeted attacks, it increases the minimum number of bit flips required to break the model by 5.6× to 14.3×; under targeted attacks, it effectively reduces the attack success rate (ASR) to 1.40%–15.60%. Moreover, the method can be fully deployed at model export time, yielding extremely low runtime overhead across CNN and Transformer architectures. These results indicate that, compared with global redundancy-based defenses, KHAD combines structural propagation properties with data-driven activity to precisely identify and block a small number of high-leverage propagation paths, trading a slight cost for substantial robustness gains. Jiajun Guo, Zhezhao Yang, Tianqi Shi, Jie Xiao 0003 |
TrustCom | 5 |
| 2025 | Pruning for Security: Mitigating Stealthy Bit-Flip Attacks with Efficiency GainsabstractStealthy Bit-Flip Attacks, which manipulate DNN predictions by altering a few hardware-level weights, pose a severe threat to safety-critical applications. To address this threat, we propose the Dual-Mask Dynamic Defense (DMDD), whose core idea is to proactively and dynamically disrupt the static paths that attackers construct by exploiting model redundancy at inference time. DMDD’s defense is accomplished by two synergistic online mechanisms: 1) At the intermediate layers, Dual-mask Dynamic Neuron Pruning (DDNP) first employs an "Efficiency Mask" to prune task-irrelevant redundant neurons for efficiency gains, and then uses a "Security Mask" to sever potential attack paths by monitoring statistical anomalies in neuron behavior. 2) At the final layer, Dynamic Connection Purification (DCP) is activated to perform semantic-level filtering. To ensure the model can tolerate this online restructuring and the effects of residual attack paths, we also designed a complementary offline Pruning-aware Robustness Training (PART) strategy to enhance the model’s pruning tolerance and enforce larger inter-class distances. Unlike traditional defenses, DMDD transforms defense overhead into performance gains, achieving security enhancement with "negative computational overhead." Experiments show that DMDD can suppress the Attack Success Rate (ASR) of mainstream attacks like TBFA and TA-LBF from nearly 100% to below 15%, while simultaneously reducing the model’s computational load (FLOPs) by 26%-42%, providing an efficient and practical solution for secure DNN deployment in resource-constrained scenarios. Jiajun Guo, Zhezhao Yang, Tianqi Shi, Jie Xiao 0003 |
TrustCom | 5 |
| 2024 | Life regression based patch slimming for vision transformers
Tianqi Shi, Lechao Cheng, Zunlei Feng, Mingli Song |
Neural Networks | 4 |
| 2023 | How to Prevent the Continuous Damage of Noises to Model Training?abstractDeep learning with noisy labels is challenging and inevitable in many circumstances. Existing methods reduce the impact of mislabeled samples by reducing loss weights or screening, which highly rely on the model's superior discriminative power for identifying mislabeled samples. However, in the training stage, the trainee model is imperfect and will wrongly predict some mislabeled samples, which cause continuous damage to the model training. Consequently, there is a large performance gap between existing anti-noise models trained with noisy samples and models trained with clean samples. In this paper, we put forward a Gradient Switching Strategy (GSS) to prevent the continuous damage of mislabeled samples to the classifier. Theoretical analysis shows that the damage comes from the misleading gradient direction computed from the mislabeled samples. The trainee model will deviate from the correct optimization direction under the influence of the accumulated misleading gradient of mislabeled samples. To address this problem, the proposed GSS alleviates the damage by switching the gradient direction of each sample based on the gradient direction pool, which contains all-class gradient directions with different probabilities. During training, each gradient direction pool is updated iteratively, which assigns higher probabilities to potential principal directions for high-confidence samples. Conversely, uncertain samples are forced to explore in different directions rather than mislead model in a fixed direction. Extensive experiments show that GSS can achieve comparable performance with a model trained with clean data. Moreover, the proposed GSS is pluggable for existing frameworks. This idea of switching gradient directions provides a new perspective for future noisy-label learning. Xiaotian Yu, Tianqi Shi, Zunlei Feng, Mingli Song |
CVPR | 3 |
| 2023 | A Loopback Network for Explainable Microvascular Invasion ClassificationabstractMicrovascular invasion (MVI) is a critical factor for prognosis evaluation and cancer treatment. The current diagnosis of MVI relies on pathologists to manually find out cancerous cells from hundreds of blood vessels, which is time-consuming, tedious, and subjective. Recently, deep learning has achieved promising results in medical image analysis tasks. However, the unexplainability of black box models and the requirement of massive annotated samples limit the clinical application of deep learning based diagnostic methods. In this paper, aiming to develop an accurate, objective, and explainable diagnosis tool for MVI, we propose a Loopback Network (LoopNet) for classifying MVI efficiently. With the image-level category annotations of the collected Pathologic Vessel Image Dataset (PVID), LoopNet is devised to be composed binary classification branch and cell locating branch. The latter is devised to locate the area of cancerous cells, regular non-cancerous cells, and background. For healthy samples, the pseudo masks of cells supervise the cell locating branch to distinguish the area of regular non-cancerous cells and background. For each MVI sample, the cell locating branch predicts the mask of cancerous cells. Then the masked cancerous and non-cancerous areas of the same sample are input back to the binary classification branch separately. The loopback between two branches enables the category label to supervise the cell locating branch to learn the locating ability for cancerous areas. Experiment results show that the proposed LoopNet achieves 97.5% accuracy on MVI classification. Surprisingly, the proposed loopback mechanism not only enables LoopNet to predict the cancerous area but also facilitates the classification backbone to achieve better classification performance. Shengxuming Zhang, Tianqi Shi, Xiuming Zhang, Jie Lei 0002, Zunlei Feng, Mingli Song |
CVPR | 2 |
| 2023 | Spectral Energy Model-Driven Inversion of XCO2 in IPDA Lidar Remote SensingabstractCarbon observation satellites based on passive theory (e.g., OCO-2/3, GOSAT-1/2, and TanSat) have relatively high carbon dioxide column concentration (XCO2) accuracy when the observation conditions are met. Passive satellites have data bias and coverage deficiencies due to cloud cover, low albedo, low-light conditions, and aerosol scattering, resulting in carbon observation satellites based on passive theory that cannot meet the demand for high-precision, all-day, all-weather XCO2 monitoring. Active detection satellites are urgently needed to support global carbon sources, sinks, and carbon neutrality. China intends to launch a sensor satellite with active detection of XCO2 in the coming years. In this work, based on the satellite’s scaled-down airborne experiments, a spectral energy model was developed to optimize the conventional inversion algorithm and achieve a more accurate XCO2 inversion. The 1.572-$\mu \text{m}$integrated path differential absorption (IPDA) lidar column length is used indirectly to evaluate the accuracy of the spectral energy model for signal extraction. Also, the experimental results show that the accuracy of the signal extracted by the 1.572-$\mu \text{m}$IPDA lidar column length is 0.74 and 6.20 m at sea and on land based on the indirect evaluation of the length of the 1.572-$\mu \text{m}$IPDA lidar column length. The optimized XCO2 was evaluated (standard deviation as an evaluation metric) and its XCO2 standard deviation reduced by 31%, 63%, and 66% in the ocean, plains, and mountains, respectively. Our algorithm can obtain the XCO2 with a consistent trend by using XCO2 from the OCO-2 satellite as a reference. The calculated XCO2 is more accurate in areas dominated by anthropogenic factors (plains), due to the accuracy of the IPDA detection mechanism. This algorithm improves the accuracy and robustness of XCO2 inversion and has important reference significance for the IPDA lidar carried by China’s satellites to be launched in this year. Ge Han, Xin Ma 0007, Tianqi Shi, Jianye Yuan, Wanqin Zhong, Yanran Peng, Wei Gong 0004 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2022 | Potential of Ground-Based Multiwavelength Differential Absorption LiDAR to Measure δ¹³C in Open Detected PathabstractA novel framework was proposed to measure atmospheric concentration of$\delta ^{13}C$using a multiwavelength integrated path differential absorption (IPDA) LiDAR. The spectroscopy range of the multiwavelength IPDA LiDAR is recommended from 2264.5 to 2265.5 cm$^{-1}$. Using the proposed retrieving method, the relative error of$\delta ^{13}C$retrievals would be within 0.16‰ under reasonable settings. Moreover, the proposed method shows reliable performances in different circumstances. It would be of great significance for exploring the characteristic of$\delta ^{13}C$in the ecosystem and anthropogenic emissions in the future. Tianqi Shi, Ge Han, Xin Ma 0007, Wei Gong 0004, Zhipeng Pei, Ruonan Qiu |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2022 | Improving CO₂ Concentration Profile Measurements From a Ground-Based CO₂-DIAL Through Conditional AdjustmentabstractGround-based differential absorption lidar (DIAL) can measure vertical CO2concentration profiles in the troposphere. Here, we propose a method of improving the accuracy and precision of CO2concentration profiles measurements. This method combines a conditional adjustment with Chebyshev fitting to reduce the error of the retrieved results in view of the received signal around the atmospheric boundary layer (ABL) with a high signal-to-noise ratio (SNR). Simulation experiments verified the effectiveness of this method. The accuracy of CO2concentration profiles can be improved larger than 83.4% when compared with that via traditional methods, and the standard deviation of the measured CO2concentration profiles calculated by our method was reduced by approximately 0.43–22.51 ppm when compared with the results calculated by traditional methods. Two real cases in different locations were also examined with the proposed technique. The results indicated the applicability of our method in measuring other trace gases by using DIAL. Tianqi Shi, Xin Ma 0007, Ge Han, Zhipeng Pei, Wei Gong 0004 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2022 | An improved indoor pedestrian dead reckoning algorithm using ambient light and sensors
Xiaoxiao Tao, Tianqi Shi, Xin Ma 0007, Zhipeng Pei |
Multim. Tools Appl. | 2 |
| 2022 | A Method for Estimating the Background Column Concentration of CO2 Using the Lagrangian ApproachabstractWith the rapid growth of GHG monitoring satellites, more and more studies focused on the issue of inversion/optimization of CO2 fluxes using satellite-derived XCO2 observations in recent years. A common and critical challenge in this framework is the separation of background and anomalies from XCO2 observations, which directly affect performance of the CO2 fluxes inversion. We proposed a novel method to accurately extract background XCO2 from satellite observations. A series of observing system simulation experiments were performed to test the performance of the method. We found that the bias and uncertainty of the background concentration are below 0.01 ppm and 0.05 ppm in the given cases, respectively. Based on this method, we selected five overpasses from 2014 to 2016 to demonstrate a regional-scale flux inversion near Riyadh. The comparison with the two previous methods shows that the posterior simulated XCO2 by the method proposed in this paper can match better with the observed XCO2 from OCO-2. Zhipeng Pei, Ge Han, Xin Ma 0007, Tianqi Shi, Wei Gong 0004 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2021 | Measuring Co2 Concentration by Airborne LidarabstractCO2 is the most important warming gas in atmosphere, it's meaningful to measure the CO2 concentration with high precise by different sensors. In this manuscript, we introduced the campaign of QHD (Qinhuangdao) flights, which equipped with CO2- IPDA (integrated path of different absorption LIDAR), an in-situ CO2 sensor, thermometer, hygrometer and GPS (Global Positioning System). A fast and accurately retrieve method of CO2 by the oral data acquired by the airborne IPDA has been developed by our group. These flight contains three different landforms, including sea, city and mountain. It shows apparent difference on the distribution of CO2 among different landforms. Finally, we compared XCO2 data from OCO-2 and results of XCO2 calculated by IPDA system, it shows little difference. Tianqi Shi, Ge Han, Xin Ma 0007 |
IGARSS | 1 |
| 2021 | A Regional Spatiotemporal Downscaling Method for CO2 ColumnsabstractQuantification of the distribution of the CO2dry-air mixing ratio (XCO2) is crucial for understanding the carbon cycle. However, clouds and aerosols in the line of light create spectral interference with CO2signals. This interference can result in a low yield of XCO2retrievals, thus limiting the application of these valuable satellite data. In this study, we developed an innovative methodology to obtain XCO2maps of high spatial and temporal resolution using satellite data. The method first interpolates the spatial properties using an empirical Bayesian kriging (EBK) algorithm. Then, the temporal properties are modulated based on a CO2curve database that was constructed using temporal contours and transfer learning techniques. We applied this method to obtain spatiotemporal XCO2maps over mainland China using the Orbiting Carbon Observatory 2 (OCO-2) data product OCO-2_L2_Lite_FP 9r for the period from January 1 to December 31, 2019. The correlation coefficient ($R^{2}$) was 0.8056, and the average absolute prediction error [root-mean-square error (RMSE)] was 0.9951. In the research area of mainland China, the vacancy validation strategy was adopted and yielded$R^{2}$and RMSE of 0.8230 and 0.9746, respectively. We used the 2018–2019 ground-based data from four Total Carbon Column Observing Network (TCCON) sites in Europe and 2016 Hefei sites in mainland China to evaluate the performance of this new mapping method, respectively. Also, we obtained$R^{2}$of 0.8690 and the RMSE of 0.9056 in Europe and$R^{2}$of 0.8473 and the RMSE of 0.7026 in mainland China, proving the robustness and high precision of our method. This mapping technique is capable of filling the spatiotemporal gaps of satellite measurements with the high accuracy and resolution needed for its scientific application; thus, it has the potential to augment the scientific returns of satellite missions (e.g., USA OCO-2 Japan GOSAT and Chinese TanSat). Xin Ma 0007, Ge Han, Feiyue Mao, Tianqi Shi, Tongtong Sun, Wei Gong 0004 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2019 | Social MIL: Interaction-Aware for Crowd Anomaly DetectionabstractCrowd anomaly detection under surveillance scene is a quite challenging task, which often companies with not rare objects, unexpected bursts in activity and complex dynamic patterns. In this paper, we propose a social multiple-instance learning(MIL) framework with a dual-branch network by considering dynamic interaction among groups, individuals and environment to obtain attentive spatial-temporal feature representation. First, MIL is employed to overcome the challenge of rare training abnormal samples and video-based labels. The social force map is utilized for modeling behavior interaction to supply the prior knowledge. In addition, we introduce the self-attention module, which represents a more discriminative spatial-temporal feature based on C3D network through implementing weight redistribution inside the feature. The results of the experiments conducted on UCF-Crime dataset show that the proposed dual-branch social multiple-instance learning (MIL) anomaly detection framework with the dual-branch network outperforms than existing approaches and obtains the state-of-the-art performance. Shuheng Lin, Hua Yang 0001, Xianchao Tang, Tianqi Shi, Lin Chen 0019 |
AVSS | 4 |