EDBT 2026 Demo / reviewers in the wild / expert
Wenxuan Xie
dblp:142/0064
· DBLP profile ↗
34ranked-venue papers
8as first author
26since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 4 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 4 first-author · 7 since 2021Computer networks · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ReEx-SQL: Reasoning with Execution-Aware Reinforcement Learning for Text-to-SQLabstractYaxun Dai, Wenxuan Xie, Xialie Zhuang, Tianyu Yang, Ziyi Liu, Haiqin Yang, Yiying Yang, Yuhang Zhao, Pingfu Chao, Wenhao Jiang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yaxun Dai, Wenxuan Xie, Xialie Zhuang, Tianyu Yang 0003, Ziyi Liu 0005, Haiqin Yang, Pingfu Chao |
ACL (1) | 2 |
| 2026 | SDE-SQL: Enhancing Text-to-SQL Generation in Large Language Models via Self-Driven Exploration with SQL ProbesabstractRecent advances in large language models (LLMs) have led to substantial progress on the Text-to-SQL task.However, existing approaches typically depend on static, preprocessed database information supplied at inference time, which restricts the model's capacity to deeply comprehend the underlying database content.In the absence of dynamic interaction, LLMs are limited to fixed, humancurated context and lack the ability to autonomously query or explore the data.To overcome this limitation, we introduce SDE-SQL, a novel framework that empowers LLMs to perform Self-Driven Exploration of databases during inference.This is achieved through the generation and execution of SQL probes, enabling the model to actively retrieve information and iteratively refine its understanding of the database.Unlike prior methods, SDE-SQL operates in a zero-shot setting, requiring no in-context demonstrations or question-SQL pairs.Evaluated on the BIRD benchmark with Qwen2.5-72B-Instruct,SDE-SQL achieves an 8.02% relative improvement in execution accuracy over the vanilla Qwen2.5-72B-Instructbaseline, establishing a new state-of-the-art among open-source methods without supervised fine-tuning (SFT) or model ensembling.Furthermore, when combined with SFT, SDE-SQL delivers an additional 0.52% performance gain.Our code is publicly available at https://github.com/ Lancelot-Xie/SDE-SQL. Wenxuan Xie, Yaxun Dai |
ACL (1) | 1 |
| 2026 | Long Video Understanding With Learnable Retrieval in Video-Language ModelsabstractThe remarkable natural language understanding, reasoning, and generation capabilities of large language models (LLMs) have made them attractive for application to video understanding, utilizing video tokens as contextual input. However, employing LLMs for long video understanding presents significant challenges. The extensive number of video tokens leads to considerable computational costs for LLMs while using aggregated tokens results in loss of vision details. Moreover, the presence of abundant question-irrelevant tokens introduces noise to the video reasoning process. To address these issues, we introduce a simple yet effective learnable retrieval-based video-language model (R-VLM) for efficient long video understanding. Specifically, given a question and a long video, our model identifies the most relevant$K$video chunks and uses their associated visual tokens to serve as context for the LLM inference. This effectively reduces the number of video tokens, eliminates noise interference, and enhances system performance. We achieve this by incorporating a learnable lightweight MLP block to facilitate the efficient retrieval of question-relevant chunks, through the end-to-end training of our video-language model with a proposed soft matching loss. Experimental results on multiple zero-shot video question answering datasets validate the effectiveness of our framework for comprehending long videos. Cuiling Lan, Wenxuan Xie, Xuejin Chen, Yan Lu 0001 |
IEEE Trans. Multim. | 3 |
| 2025 | Towards Practical Real-Time Neural Video CompressionabstractWe introduce a practical real-time neural video codec (NVC) designed to deliver high compression ratio, low latency and broad versatility. In practice, the coding speed of NVCs depends on 1) computational costs, and 2) non-computational operational costs, such as memory I/O and the number of function calls. While most efficient NVCs prioritize reducing computational cost, we identify operational cost as the primary bottleneck to achieving higher coding speed. Leveraging this insight, we introduce a set of efficiency-driven design improvements focused on minimizing operational costs. Specifically, we employ implicit temporal modeling to eliminate complex explicit motion modules, and use single low-resolution latent representations rather than progressive downsampling. These innovations significantly accelerate NVC without sacrificing compression quality. Additionally, we implement model integerization for consistent cross-device coding and a module-bank-based rate control scheme to improve practical adaptability. Experiments show our proposed DCVC-RT achieves an impressive average encoding/decoding speed at 125.2/112.8 fps (frames per second) for 1080p video, while saving an average of 21% in bitrate compared to H.266/VTM. The code is available at https://github.com/microsoft/DCVC. Zhaoyang Jia, Bin Li 0012, Jiahao Li 0001, Wenxuan Xie, Houqiang Li, Yan Lu 0001 |
CVPR | 4 |
| 2025 | Self-Sufficient 5-DoF Discrete Global Localization for Magnetically-Actuated Endoscope in BronchoscopyabstractExisting sensor-based global localization methods limit the miniaturization potential of magnetically-actuated endoscopes (MAE) while localization based on external medical imaging demands accurate registration and imposes a variety of modality-specific challenges during continuous image acquisition. This work proposes a novel self-sufficient method for discrete (one-time) global localization of an MAE based solely on inherent endoscopic images without any prior MAE pose information. More specifically, it adopts a model-free control approach to determine five different external magnet (EM) poses (corresponding to five independent nonlinear equations) that can align the MAE image center with the lumen center while the MAE maintains the same pose. The five degree-of-freedom (DoF) global pose of the MAE can then be estimated by minimizing the root mean square of MAE's torque balance residuals under these EM poses. Our proposed method achieves similar accuracy as other sensor-based methods for permanent magnet-driven MAE with$\mathbf{6.7} \pm \mathbf{2.1}$mm position error and$\mathbf{9.5} \pm \mathbf{2.9}^{\circ}$orientation error in the experiments. Compared to existing methods, our approach does not require physical sensor integration, enabling a more compact endoscope design for exploration in narrower respiratory tracts. It also offers a critical step toward achieving sensorless and continuous global localization of the permanent magnet-driven MAE during its autonomous navigation. Jiewen Tan, Wenxuan Xie, Shing Shin Cheng |
ICRA | 4 |
| 2025 | Motion-Guided Dual-Camera Tracker for Endoscope Tracking and Motion Analysis in a Mechanical Gastric SimulatorabstractFlexible endoscope motion tracking and analysis in mechanical simulators have proven useful for endoscopy training. Common motion tracking methods based on electromagnetic tracker are however limited by their high cost and material susceptibility. In this work, the motion-guided dual-camera vision tracker is proposed to provide robust and accurate tracking of the endoscope tip's 3D position. The tracker addresses several unique challenges of tracking flexible endoscope tip inside a dynamic, life-sized mechanical simulator. To address the appearance variation and keep dualcamera tracking consistency, the cross-camera mutual template strategy (CMT) is proposed by introducing dynamic transient mutual templates. To alleviate large occlusion and light-induced distortion, the Mamba-based motion-guided prediction head (MMH) is presented to aggregate historical motion with visual tracking. The proposed tracker achieves superior performance against state-of-the-art vision trackers, achieving 42% and 72% improvements against the second-best method in average error and maximum error. Further motion analysis involving novice and expert endoscopists also shows that the tip 3D motion provided by the proposed tracker enables more reliable motion analysis and more substantial differentiation between different expertise levels, compared with other trackers. Project page: https://github.com/PieceZhang/MotionDCTrack Yuelin Zhang, Kim Yan, Chun Ping Lam, Chengyu Fang 0001, Wenxuan Xie, Yufu Qiu, Raymond Shing-Yan Tang, Shing Shin Cheng |
ICRA | 5 |
| 2025 | CIDD: Collaborative Intelligence for Structure-Based Drug Design Empowered by LLMsabstractStructure-guided molecular generation is pivotal in early-stage drug discovery, enabling the design of compounds tailored to specific protein targets. However, despite recent advances in 3D generative modeling, particularly in improving docking scores, these methods often produce rare and intrinsically irrational molecular structures that deviate from drug-like chemical space. To quantify this issue, we propose a novel metric, the Molecule Reasonable Ratio (MRR), which measures structural rationality and reveals a critical gap between existing models and real-world approved drugs. To address this, we introduce the Collaborative Intelligence Drug Design (CIDD) framework, the first approach to unify the 3D interaction modeling capabilities of generative models with the general knowledge and reasoning power of large language models (LLMs). By leveraging LLM-based Chain-of-Thought reasoning, CIDD generates molecules that not only bind effectively to protein pockets but also exhibit strong structural drug-likeness, rationality, and synthetic accessibility. On the CrossDocked2020 benchmark, CIDD consistently improves drug-likeness metrics, including QED, SA, and MRR, across different base generative models, while maintaining competitive binding affinity. Notably, it raises the combined success rate (balancing drug-likeness and binding) from 15.72% to 34.59%, more than doubling previous results. These findings demonstrate the value of integrating knowledge reasoning with geometric generation to advance AI-driven drug design. Yanwen Huang, Yiqiao Liu, Wenxuan Xie, Bowei He, Haichuan Tan, Wei-Ying Ma, Ya-Qin Zhang, Yanyan Lan |
NeurIPS | 4 |
| 2025 | Uncertain Free Disposal Hull Model with Application to Chinese BanksabstractAs an alternative model to the data envelopment analysis (DEA), the free disposal hull (FDH) model has excellent performance in measuring decision making unit (DMU) efficiency under the condition that the production possibilities do not satisfy the convexity assumption. However, in actual production and life, many data are difficult to collect and cannot obtain accurate values, such as carbon dioxide emissions. In the case of imprecise data, the FDH model cannot evaluate the efficiency of DMU. In this context, a new uncertain FDH model based on uncertainty theory is proposed. In this paper, the uncertain FDH model is solved precisely by means of the uncertain chance constraint method and the expected value method. Finally, an application example of measuring the efficiency of 30 banks in China in 2019 is given to documenting the feasibility of the proposed model. Jiali Wu, Wenxuan Xie, Yuhong Sheng |
Int. J. Uncertain. Fuzziness Knowl. Based Syst. | 2 |
| 2025 | UAV Identification via Multiscale Decision Fusion CNN Utilizing Micro-Doppler FeaturesabstractWith the rapid expansion of the low-altitude economy, the supervision of low-altitude unmanned aerial vehicles (UAVs) is encountering increasingly complex challenges, with the accurate identification of UAVs emerging as a critical issue. Radar systems, owing to their robustness against external interference, are frequently integrated with other technologies to enhance UAV identification capabilities. This study introduces a novel UAV radar signal identification approach utilizing a multi-scale residual convolutional neural network (CNN). By combining time-domain features, time-frequency domain features, and range-Doppler features from a frequency-modulated continuous-wave radar (FMCWR), and employing decision-level fusion techniques to extract multi-scale characteristics, the proposed method significantly enhances feature representation. Experimental results conclusively demonstrate that this fusion strategy achieves a identification accuracy of 94%. Hanchu Zhou, Yijun Chen 0001, Anmin Gong, Zhu Yongzhong, Caijing Mo, Wenxuan Xie |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2025 | Deep Reinforcement Learning-Based Computation Computational Offloading for Space-Air-Ground Integrated Vehicle NetworksabstractIn remote or disaster areas, where terrestrial networks are difficult to cover and Terrestrial Edge Computing (TEC) infrastructures are unavailable, solving the computation computational offloading for Internet of Vehicles (IoV) scenarios is challenging. Current terrestrial networks have high data rates, great connectivity, and low delay, but global coverage is limited. Space–Air–Ground Integrated Networks (SAGIN) can improve the coverage limitations of terrestrial networks and enhance disaster resistance. However, the rising complexity and heterogeneity of networks make it difficult to find a robust and intelligent computational offload strategy. Therefore, joint scheduling of space, air, and ground resources is needed to meet the growing demand for services. In light of this, we propose an integrated network framework for Space-Air Auxiliary Vehicle Computation (SA-AVC) and build a system model to support various IoV services in remote areas. Our model aims to maximize delay and fair utility and increase the utilization of satellites and Autonomous aerial vehicles (AAVs). To this end, we propose a Deep Reinforcement Learning algorithm to achieve real-time computational computational offloading decisions. We utilize the Rank-based Prioritization method in Prioritized Experience Replay (PER) to optimize our algorithm. We designed simulation experiments for validation and the results show that our proposed algorithm reduces the average system delay by 17.84%, 58.09%, and 58.32%, and the average variance of the task completion delay will be reduced by 29.41%, 48.74%, and 49.58% compared to the Deep Q Network (DQN), Q-learning and RandomChoose algorithms. Wenxuan Xie, Chen Chen 0006, Ying Ju 0001, Jun Shen 0001, Qingqi Pei, Houbing Song |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2024 | Text Grouping Adapter: Adapting Pre-Trained Text Detector for Layout AnalysisabstractSignificant progress has been made in scene text detection models since the rise of deep learning, but scene text layout analysis, which aims to group detected text instances as paragraphs, has not kept pace. Previous works either treated text detection and grouping using separate models, or train a model from scratch while using a unified one. All of them have not yet made full use of the already well-trained text detectors and easily obtainable detection datasets. In this paper, we present Text Grouping Adapter (TGA), a module that can enable the utilization of various pretrained text detectors to learn layout analysis, allowing us to adopt a well-trained text detector right off the shelf or just fine-tune it efficiently. Designed to be compatible with various text detector architectures, TGA takes detected text regions and image features as universal inputs to as-semble text instance features. To capture broader contextual information for layout analysis, we propose to predict text group masks from text instance features by one-to-many assignment. Our comprehensive experiments demonstrate that, even with frozen pretrained models, incorporating our TGA into various pretrained text detectors and text spotters can achieve superior layout analysis performance, simultaneously inheriting generalized text detection ability from pretraining. In the case of full parameter fine-tuning, we can further improve layout analysis performance. Tianci Bi, Zhizheng Zhang 0004, Wenxuan Xie, Cuiling Lan, Yan Lu 0001, Nanning Zheng 0001 |
CVPR | 4 |
| 2024 | A Topology Reconfiguration Algorithm for Cross-Domain Multi-System Measurement and Control Network Based on Minimum ConnectivityabstractMulti-network integration represents the future trend in measurement, control and communication. To provide optimal service, it is significant to study the network topology strategy of heterogeneous nodes. However, the mobility and node damage and failures caused by interference pose a huge challenge in reconfiguring the topology to ensure network performance. To address this issue, this paper proposes a topology reconfiguration algorithm based on minimum connectivity for a cross-domain multi-system measurement and control network, which integrates air, ground, and sea. This algorithm designs an optimal topology reconfiguration strategy considering connectivity, energy consumption, and position accuracy, and constructs a reliable measurement and control system. Meanwhile, by implementing the improved simulated annealing optimization, a robust measurement and control network topology can be obtained. Theoretical analysis and simulation results demonstrate that this algorithm can effectively resist external interference and ensure positioning accuracy, proved to be feasible and effective compared to the Global K-connected Energy-aware Topology-control Algorithm (GKETA) and greedy algorithm. Huilin Wang, Lidong Zhu, Wenxuan Xie |
ISNCC | 4 |
| 2024 | A Non-Terrestrial Network Congestion Control Scheme Based on the SAC MethodabstractSatellite Internet technology is currently booming. Due to the high cost of satellites and the limited number of orbits, a satellite must cover a large area and serve more users than before. Given that the random access process is the initial step in establishing communication between satellites and terrestrial users, congestion control has emerged as a focal point of research. In this paper, we address the random access congestion control problem in satellite Internet scenarios and introduce the soft actor-critic (SAC) method to dynamically set the access class barring (ACB) factors. The throughput of random access is optimized to allow a single satellite to support a greater number of users without experiencing congestion. The SAC-ACB method is verified by simulation to converge faster and perform better than the proximal policy optimization (PPO) method. Wenxuan Xie, Lidong Zhu, Yanjun Song, Huilin Wang |
ISNCC | 1 |
| 2024 | Slot-VLM: Object-Event Slots for Video-Language ModelingabstractVideo-Language Models (VLMs), powered by the advancements in Large Language Models (LLMs), are charting new frontiers in video understanding. A pivotal challenge is the development of an effective method to encapsulate video content into a set of representative tokens to align with LLMs. In this work, we introduce Slot-VLM, a new framework designed to generate semantically decomposed video tokens, in terms of object-wise and event-wise visual representations, to facilitate LLM inference. Particularly, we design an Object-Event Slots module, i.e., OE-Slots, that adaptively aggregates the dense video tokens from the vision encoder to a set of representative slots. In order to take into account both the spatial object details and the varied temporal dynamics, we build OE-Slots with two branches: the Object-Slots branch and the Event-Slots branch. The Object-Slots branch focuses on extracting object-centric slots from features of high spatial resolution but low frame sample rate, emphasizing detailed object information. The Event-Slots branch is engineered to learn event-centric slots from high temporal sample rate but low spatial resolution features. These complementary slots are combined to form the vision context, serving as the input to the LLM for effective video reasoning. Our experimental results demonstrate the effectiveness of our Slot-VLM, which achieves the state-of-the-art performance on video question-answering. Cuiling Lan, Wenxuan Xie, Xuejin Chen, Yan Lu 0001 |
NeurIPS | 3 |
| 2024 | CRL-MABA: A Completion Rate Learning-Based Accurate Data Collection Scheme in Large-Scale Energy InternetabstractThe Energy Internet (EI) aims to build a sustainable energy ecosystem by connecting diverse energy sources and prosumers. Mobile Crowd Sensing (MCS) enables efficient data collection for monitoring and aggregation from distributed devices. Given the complex behavior of workers driven by self-interest, recruiting trustworthy, high-quality, and inexpensive workers remains a significant challenge in research and practice. Previous studies often assume that worker characteristics are known or can be obtained after data collection. However, evaluating worker qualities is quite challenging in the face of multi-source data and complex workers. To address this, we propose a Completion Rate Learning based Multi-Armed Bandit reverse Auction (CRL-MABA) scheme for identifying and selecting high-quality workers in MCS. Our CRL-MABA scheme first proposes a Spatial-Temporal Upper Confidence Bound (STUCB) method to recruit workers, considering both the quality of workers for exploitation and the spatiotemporal features for exploration. In addition, the Dual-Stage Data Estimation Mechanism (DSDEM) and Long-Term and Short-Term Memory Learning (LTSTML) are designed to identify workers accurately and efficiently. Importantly, our proposed scheme avoids the impractical assumptions in previous works while satisfying important criteria such as truthfulness, individual rationality, and computational efficiency. The effectiveness of our scheme is demonstrated through extensive experimental results, which show its superiority over existing strategies. Kejia Fan, Jianheng Tang 0001, Wenxuan Xie, Feijiang Han, Yajiang Huang, Zhenzhe Qu, Anfeng Liu, Naixue Xiong, Tian Wang 0001, Shaobo Zhang 0001 |
IEEE Internet Things J. | 3 |
| 2024 | BTV-CMAB: A Bi-Directional Trust Verification-Based Combinatorial Multiarmed Bandit Scheme for Mobile CrowdsourcingabstractMobile crowdsourcing (MCS) is an emerging paradigm that harnesses the collective power of the crowd to tackle large-scale tasks. To ensure the high-quality worker selection, various combinatorial multiarmed bandit (CMAB)-based schemes have been proposed. However, previous schemes often overlook critical issues. First, the post-unknown worker recruitment (PUWR) problem emerges when the quality of a worker remains unknown despite reported worker data. Second, the presence of Sybil Requesters is often neglected, who manipulate ratings to deceive workers for malicious purposes. To tackle these challenges, we present an innovative scheme called bi-directional trust verification-based CMAB (BTV-CMAB). First, we propose a truth quality discovery approach that effectively addresses the PUWR problem by estimating worker quality. Additionally, we employ a BTV mechanism to assess the Degree of Trust (DoT) of requesters and the reputation of workers. To select top-notch workers for MCS, we combine the worker quality and reputation into an upper confidence bound (UCB) index. The effectiveness of the BTV-CMAB scheme is supported by theoretical proof, which demonstrates its ability to ensure truthfulness and individual rationality. Furthermore, experimental results reveal promising improvements achieved by our scheme, including a 17.44%, increase in the platform’s revenue and a significant decrease in regret of up to 88.26%. To the best of our knowledge, this study is the first to propose utilizing a BTV mechanism to effectively address the PUWR problem and counter the threat of Sybil attacks in the CMAB-based worker recruitment process. Jianheng Tang 0001, Kejia Fan, Wenxuan Xie, Feijiang Han, Zhenzhe Qu, Anfeng Liu, Naixue Xiong, Shaobo Zhang 0001, Tian Wang 0001 |
IEEE Internet Things J. | 3 |
| 2024 | Uniform gradient magnetic field and spatial localization method based on Maxwell coils for virtual surgery simulationabstractAbstract With the development of virtual reality technology, simulation surgery has become a low‐risk surgical training method and high‐precision positioning of surgical instruments is required in virtual simulation surgery. In this paper we design and validate a novel electromagnetic positioning method based on a uniform gradient magnetic field. We employ Maxwell coils to generate the uniform gradient magnetic field and propose two positioning algorithms based on magnetic field, namely the linear equation positioning algorithm and the magnetic field fingerprint positioning algorithm. After validating the feasibility of proposed positioning system through simulation, we construct a prototype system and conduct practical experiments. The experimental results demonstrate that the positioning system exhibits excellent accuracy and speed in both simulation and real‐world applications. The positioning accuracy remains consistent and high, showing no significant variation with changes in the positions of surgical instruments. Xutian Deng, Xujie Zhao, Wenxuan Xie, Jianhui Zhao 0001 |
Comput. Animat. Virtual Worlds | 4 |
| 2023 | Unifying Layout Generation with a Decoupled Diffusion ModelabstractLayout generation aims to synthesize realistic graphic scenes consisting of elements with different attributes in-cluding category, size, position, and between-element relation. It is a crucial task for reducing the burden on heavyduty graphic design works for formatted scenes, e.g., publications, documents, and user interfaces (UIs). Diverse application scenarios impose a big challenge in unifying various layout generation subtasks, including conditional and unconditional generation. In this paper, we propose a Layout Diffusion Generative Model (LDGM) to achieve such unification with a single decoupled diffusion model. LDGM views a layout of arbitrary missing or coarse element attributes as an intermediate diffusion status from a completed layout. Since different attributes have their individual semantics and characteristics, we propose to decouple the diffusion processes for them to improve the diversity of training samples and learn the reverse process jointly to exploit global-scope contexts for facilitating generation. As a result, our LDGM can generate layouts either from scratch or conditional on arbitrary available attributes. Extensive qualitative and quantitative experiments demonstrate our proposed LDGM outperforms existing layout generation models in both functionality and performance. Mude Hui, Zhizheng Zhang 0004, Wenxuan Xie, Yuwang Wang, Yan Lu 0001 |
CVPR | 4 |
| 2023 | A Semi-supervised Sensing Rate Learning based CMAB scheme to combat COVID-19 by trustful data collection in the crowd
Jianheng Tang 0001, Kejia Fan, Wenxuan Xie, Luomin Zeng, Feijiang Han, Guosheng Huang, Tian Wang 0001, Anfeng Liu, Shaobo Zhang 0001 |
Comput. Commun. | 3 |
| 2023 | Credit and quality intelligent learning based multi-armed bandit scheme for unknown worker selection in multimedia MCS
Jianheng Tang 0001, Feijiang Han, Kejia Fan, Wenxuan Xie, Pengzhi Yin, Zhenzhe Qu, Anfeng Liu, Naixue Xiong, Shaobo Zhang 0001, Tian Wang 0001 |
Inf. Sci. | 4 |
| 2023 | Parameter Extraction of Accelerated Motion Targets Based on Vortex Electromagnetic Wave RadarabstractVortex electromagnetic (EM) wave radar can obtain more accurate rotation parameters for target classification and recognition. However, the existing rotation parameter extraction methods are difficult to obtain the acceleration and initial phase of composite moving target in the case of multiple scattering points. In this paper, the estimation error of accelerated motion target features in a situation with numerous scattering points is examined for the first time, along with a method for extracting accelerated motion target parameters based on vortex electromagnetic wave radar. Firstly, the relationship between the characteristics of non derivable points (NDPs) in the rotating Doppler frequency shift and the characteristics of the target’s spin motion is analyzed, the method of estimating the number of scattering points is studied, and the formula for extracting the spin parameters is derived. After that, the time-frequency analysis of the echo signal is combined to extract the target translational motion characteristics. Simulation results show that the proposed method is effective. Lingling Zhang 0011, Zhu Yongzhong, Yijun Chen 0001, Wenxuan Xie |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | Sparse MLP for Image Recognition: Is Self-Attention Really Necessary?abstractTransformers have sprung up in the field of computer vision. In this work, we explore whether the core self-attention module in Transformer is the key to achieving excellent performance in image recognition. To this end, we build an attention-free network called sMLPNet based on the existing MLP-based vision models. Specifically, we replace the MLP module in the token-mixing step with a novel sparse MLP (sMLP) module. For 2D image tokens, sMLP applies 1D MLP along the axial directions and the parameters are shared among rows or columns. By sparse connection and weight sharing, sMLP module significantly reduces the number of model parameters and computational complexity, avoiding the common over-fitting problem that plagues the performance of MLP-like models. When only trained on the ImageNet-1K dataset, the proposed sMLPNet achieves 81.9% top-1 accuracy with only 24M parameters, which is much better than most CNNs and vision Transformers under the same model size constraint. When scaling up to 66M parameters, sMLPNet achieves 83.4% top-1 accuracy, which is on par with the state-of-the-art Swin Transformer. The success of sMLPNet suggests that the self-attention mechanism is not necessarily a silver bullet in computer vision. The code and models are publicly available at https://github.com/microsoft/SPACH. Chuanxin Tang, Guangting Wang, Chong Luo 0001, Wenxuan Xie, Wenjun Zeng 0001 |
AAAI | 5 |
| 2022 | Deep Reinforcement Learning for Load Balancing of Edge Servers in IoV
Wenxuan Xie, Chen Chen 0006, Shaohua Wan 0001 |
Mob. Networks Appl. | 2 |
| 2022 | Correction to: Deep Reinforcement Learning for Load Balancing of Edge Servers in IoV
Wenxuan Xie, Chen Chen 0006, Shaohua Wan 0001 |
Mob. Networks Appl. | 2 |
| 2021 | Unsupervised Visual Representation Learning by Tracking Patches in VideoabstractInspired by the fact that human eyes continue to develop tracking ability in early and middle childhood, we propose to use tracking as a proxy task for a computer vision system to learn the visual representations. Modelled on the Catch game played by the children, we design a Catch-the-Patch (CtP) game for a 3D-CNN model to learn visual representations that would help with video-related tasks. In the proposed pretraining framework, we cut an image patch from a given video and let it scale and move according to a pre-set trajectory. The proxy task is to estimate the position and size of the image patch in a sequence of video frames, given only the target bounding box in the first frame. We discover that using multiple image patches simultaneously brings clear benefits. We further increase the difficulty of the game by randomly making patches invisible. Extensive experiments on mainstream benchmarks demonstrate the superior performance of CtP against other video pretraining methods. In addition, CtP-pretrained features are less sensitive to domain gaps than those trained by a supervised action recognition task. When both trained on Kinetics-400, we are pleasantly surprised to find that CtP-pretrained representation achieves much higher action classification accuracy than its fully supervised counterpart on Something-Something dataset. Guangting Wang, Yizhou Zhou, Chong Luo 0001, Wenxuan Xie, Wenjun Zeng 0001, Zhiwei Xiong |
CVPR | 4 |
| 2021 | Convolutional Neural Networks for forecasting flood process in Internet-of-Things enabled smart city
Chen Chen 0006, Qiang Hui, Wenxuan Xie, Shaohua Wan 0001, Yang Zhou 0032, Qingqi Pei |
Comput. Networks | 3 |
| 2020 | Joint Time-Frequency and Time Domain Learning for Speech EnhancementabstractFor single-channel speech enhancement, both time-domain and time-frequency-domain methods have their respective pros and cons. In this paper, we present a cross-domain framework named TFT-Net, which takes time-frequency spectrogram as input and produces time-domain waveform as output. Such a framework takes advantage of the knowledge we have about spectrogram and avoids some of the drawbacks that T-F-domain methods have been suffering from. In TFT-Net, we design an innovative dual-path attention block (DAB) to fully exploit correlations along the time and frequency axes. We further discover that a sample-independent DAB (SDAB) achieves a good tradeoff between enhanced speech quality and complexity. Ablation studies show that both the cross-domain design and the SDAB block bring large performance gain. When logarithmic MSE is used as the training criteria, TFT-Net achieves the highest SDR and SSNR among state-of-the-art methods on two major speech enhancement benchmarks. Chuanxin Tang, Chong Luo 0001, Zhiyuan Zhao 0001, Wenxuan Xie, Wenjun Zeng 0001 |
IJCAI | 4 |
| 2019 | Detect or Track: Towards Cost-Effective Video Object Detection/TrackingabstractState-of-the-art object detectors and trackers are developing fast. Trackers are in general more efficient than detectors but bear the risk of drifting. A question is hence raised – how to improve the accuracy of video object detection/tracking by utilizing the existing detectors and trackers within a given time budget? A baseline is frame skipping – detecting every N-th frames and tracking for the frames in between. This baseline, however, is suboptimal since the detection frequency should depend on the tracking quality. To this end, we propose a scheduler network, which determines to detect or track at a certain frame, as a generalization of Siamese trackers. Although being light-weight and simple in structure, the scheduler network is more effective than the frame skipping baselines and flow-based approaches, as validated on ImageNet VID dataset in video object detection/tracking. Wenxuan Xie, Xinggang Wang, Wenjun Zeng 0001 |
AAAI | 2 |
| 2019 | Learning to Update for Object Tracking With Recurrent Meta-LearnerabstractModel update lies at the heart of object tracking. Generally, model update is formulated as an online learning problem where a target model is learned over the online training set. Our key innovation is to formulate the model update problem in the meta-learning framework and learn the online learning algorithm itself using large numbers of offline videos, i.e., learning to update. The learned updater takes as input the online training set and outputs an updated target model. As a first attempt, we design the learned updater based on recurrent neural networks (RNNs) and demonstrate its application in a template-based tracker and a correlation filter-based tracker. Our learned updater consistently improves the base trackers and runs faster than realtime on GPU while requiring small memory footprint during testing. Experiments on standard benchmarks demonstrate that our learned updater outperforms commonly used update baselines including the efficient exponential moving average (EMA)-based update and the well-designed stochastic gradient descent (SGD)-based update. Equipped with our learned updater, the template-based tracker achieves state-of-the-art performance among realtime trackers on GPU. Bi Li 0005, Wenxuan Xie, Wenjun Zeng 0001, Wenyu Liu 0001 |
IEEE Trans. Image Process. | 2 |
| 2014 | Cross-View Feature Learning for Scalable Social Image AnalysisabstractNowadays images on social networking websites (e.g., Flickr) are mostly accompanied with user-contributed tags, which help cast a new light on the conventional content-based image analysis tasks such as image classification and retrieval. In order to establish a scalable social image analysis system, two issues need to be considered: 1) Supervised learning is a futile task in modeling the enormous number of concepts in the world, whereas unsupervised approaches overcome this hurdle; 2) Algorithms are required to be both spatially and temporally efficient to handle large-scale datasets. In this paper, we propose a cross-view feature learning (CVFL) framework to handle the problem of social image analysis effectively and efficiently. Through explicitly modeling the relevance between image content and tags (which is empirically shown to be visually and semantically meaningful), CVFL yields more promising results than existing methods in the experiments. More importantly, being general and descriptive, CVFL and its variants can be readily applied to other large-scale multi-view tasks in unsupervised setting. Wenxuan Xie, Yuxin Peng 0001, Jianguo Xiao |
AAAI | 1 |
| 2014 | Semantic Graph Construction for Weakly-Supervised Image ParsingabstractWe investigate weakly-supervised image parsing, i.e., assigning class labels to image regions by using image-level labels only. Existing studies pay main attention to the formulation of the weakly-supervised learning problem, i.e., how to propagate class labels from images to regions given an affinity graph of regions. Notably, however, the affinity graph of regions, which is generally constructed in relatively simpler settings in existing methods, is of crucial importance to the parsing performance due to the fact that the weakly-supervised parsing problem cannot be solved within a single image, and that the affinity graph enables label propagation among multiple images. In order to embed more semantics into the affinity graph, we propose novel criteria by exploiting the weak supervision information carefully, and develop two graphs: L1 semantic graph and k-NN semantic graph. Experimental results demonstrate that the proposed semantic graphs not only capture more semantic relevance, but also perform significantly better than conventional graphs in image parsing. Wenxuan Xie, Yuxin Peng 0001, Jianguo Xiao |
AAAI | 1 |
| 2014 | Weakly-Supervised Image Parsing via Constructing Semantic Graphs and HypergraphsabstractIn this paper, we address the problem of weakly-supervised image parsing, whose aim is to automatically determine the class labels of image regions given image-level labels only. In the literature, existing studies pay main attention to the formulation of the weakly-supervised learning problem, i.e., how to propagate class labels from images to regions given an affinity graph of regions. Notably, however, the affinity graph of regions, which is generally constructed in relatively simpler settings in existing methods, is of crucial importance to the parsing performance due to the fact that the weakly-supervised image parsing problem cannot be handled within a single image, and that the affinity graph facilitates label propagation among multiple images. Therefore, in contrast to existing methods, we focus on how to make the affinity graph more descriptive through embedding more semantics into it. We develop two novel graphs by leveraging the weak supervision information carefully: 1) Semantic graph, which is established upon a conventional graph by utilizing the proposed weakly-supervised criteria; 2) Semantic hypergraph, which explores both intra-image and inter-image high-order semantic relevance. Experimental results on two standard datasets demonstrate that the proposed semantic graphs and hypergraphs not only capture more semantic relevance, but also perform significantly better than conventional graphs in image parsing. More remarkably, due to the complementariness among the proposed semantic graphs and hypergraphs, the combination of them shows even more promising results. Wenxuan Xie, Yuxin Peng 0001, Jianguo Xiao |
ACM Multimedia | 1 |
| 2014 | Graph-based multimodal semi-supervised image classification
Wenxuan Xie, Zhiwu Lu 0001, Yuxin Peng 0001, Jianguo Xiao |
Neurocomputing | 1 |
| 2013 | Multimodal semi-supervised image classification by combining tag refinement, graph-based learning and support vector regressionabstractWe investigate an image classification task where the training images come along with tags, but only a subset being labeled, and the goal is to predict the class label of test images without tags. This task is crucial for image search engine on photo sharing Web sites. In previous work, it is handled by first learning a multiple kernel learning classifier using both image content and tags to score unlabeled training images, and then building up a least-squares regression (LSR) model on visual features to predict the label of test images. However, there exist three important issues in the task: (1) Image tags on photo sharing Web sites tend to be inaccurate and incomplete, and thus refining them is beneficial; (2) Supervised learning with a limited number of labeled samples may be unreliable to some extent, while a graph-based semi-supervised approach can be adopted by also considering similarities of unlabeled data; (3) LSR is established upon centered visual kernel columns and breaks the symmetry of kernel matrix, whereas support vector regression can readily use the original visual kernel and thus leverage its full power. To handle the task more effectively, we propose to combine tag refinement, graph-based learning and support vector regression together. Experimental results on the PASCAL VOC'07 and MIR Flickr datasets show the superior performance of the proposed approach. Wenxuan Xie, Zhiwu Lu 0001, Yuxin Peng 0001, Jianguo Xiao |
ICIP | 1 |