VLDB 2026 Research / reviewers in the wild / expert
Yuxing Han 0001
dblp:91/7908-1
· DBLP profile ↗
64ranked-venue papers
4as first author
29since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 37 · 1 first-author · 15 since 2021Artificial intelligence and machine learning · 14 · 12 since 2021Computer networks · 11 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 8 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 since 2021Systems, architecture and hardware · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FedGRPO: Privately Optimizing Foundation Models with Group-Relative Rewards from Domain ClientsabstractOne important direction of Federated Foundation Models (FedFMs) is leveraging data from small client models to enhance the performance of a large server‑side foundation model. Existing methods based on model level or representation level knowledge transfer either require expensive local training or incur high communication costs and introduce unavoidable privacy risks. We reformulate this problem as a reinforcement learning style evaluation process and propose FedGRPO, a privacy preserving framework comprising two modules. The first module performs competence-based expert selection by building a lightweight confidence graph from auxiliary data to identify the most suitable clients for each question. The second module leverages the “Group Relative” concept from the Group Relative Policy Optimization (GRPO) framework by packaging each question together with its solution rationale into candidate policies, dispatching these policies to a selected subset of expert clients, and aggregating solely the resulting scalar reward signals via a federated group–relative loss function. By exchanging reward values instead of data or model updates, FedGRPO reduces privacy risk and communication overhead while enabling parallel evaluation across heterogeneous devices. Empirical results on diverse domain tasks demonstrate that FedGRPO achieves superior downstream accuracy and communication efficiency compared to conventional FedFMs baselines. Gongxi Zhu, Hanlin Gu, Lixin Fan, Qiang Yang 0001, Yuxing Han 0001 |
AAAI | 5 |
| 2026 | Beyond Ranking: Fine-Grained Diagnostics and Self-Improvement for MLLMsabstractMingze Xu, Zijing Zhao, Qiming Peng, Houwen Peng, Han Hu, Zhanhui Kang, Yuxing Han. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zijing Zhao 0010, Qiming Peng, Houwen Peng, Han Hu 0001, Zhanhui Kang, Yuxing Han 0001 |
ACL (1) | 7 |
| 2026 | Disentangling data distribution for optimal and communication-efficient federated learning
Hanlin Gu, Lixin Fan, Yuxing Han 0001, Qiang Yang 0001 |
Artif. Intell. | 4 |
| 2026 | Multi-scale semantics meet SE(3) geometry: A sequence-aware hierarchical VLA for long-horizon manipulation
Zhitao Wang, Jiangtao Wen, Roberto Horowitz, Yanke Wang, Yuxing Han 0001 |
Knowl. Based Syst. | 5 |
| 2026 | GStitch: Spatial-temporal fusion of 3D Gaussian splattings for scalable 3D reconstruction
Licheng Shen, Ho Ngai Chow, Mengqiu Wang, Yuxing Han 0001 |
Pattern Recognit. | 8 |
| 2025 | FedMIA: An Effective Membership Inference Attack Exploiting "All for One" Principle in Federated LearningabstractFederated Learning (FL) is a promising approach for training machine learning models on decentralized data while preserving privacy. However, privacy risks, particularly Membership Inference Attacks (MIAs), which aim to determine whether a specific data point belongs to a target client’s training set, remain a significant concern. Existing methods for implementing MIAs in FL primarily analyze updates from the target client, focusing on metrics such as loss, gradient norm, and gradient difference. However, these methods fail to leverage updates from non-target clients, potentially underutilizing available information. In this paper, we first formulate a one-tailed likelihood-ratio hypothesis test based on the likelihood of updates from non-target clients. Building upon this formulation, we introduce a three-step Membership Inference Attack (MIA) method, called FedMIA, which follows the "all for one"—leveraging updates from all clients across multiple communication rounds to enhance MIA effectiveness. Both theoretical analysis and extensive experimental results demonstrate that FedMIA outperforms existing MIAs in both classification and generative tasks. Additionally, it can be integrated as an extension to existing methods and is robust against various defense strategies, Non-IID data, and different federated structures. Our code is available in https://github.com/Liar-Mask/FedMIA. Gongxi Zhu, Hanlin Gu, Yuan Yao 0011, Lixin Fan, Yuxing Han 0001 |
CVPR | 6 |
| 2025 | SigmoidGS: To Guide Depth More EffectivelyabstractThe recent success of 3D Gaussian Splatting (3DGS) on the task of novel view synthesis has amazed every one with its photorealistic results with high training and rendering speed. This paper aims to increase the interpretability of the model in both geometric attribute and appearances by a simple yet effective method: lifting constraints on color features of each splat. This enables more accurate guidance from monocular depth prediction models and strengthen the ability of model to reconstruct scene geometry. Experiments shows that this effectively reduces indeterminacy of the reconstruction problem and allows better understanding of the scene structure and individual, especially in scenes with limited viewpoints, while maintaining high-fidelity rendering even in some of the uncovered views. Ho Ngai Chow, Licheng Shen, Mengqiu Wang, Yuxing Han 0001 |
ICASSP | 6 |
| 2025 | EP-SAM: An Edge-Detection Prompt SAM Based Efficient Framework for Ultra-Low Light Video SegmentationabstractThe Segment Anything Model (SAM) excels at generating high-quality object masks with various prompts but struggles in ultra-low light. We developed EP-SAM (Edge-Detection Prompt SAM) with a Low Light Edge-Detection Network (LLEN), offering strong robustness and lightweight performance in ultra-low light. Using edge information as prompts, EP-SAM achieves high-precision segmentation in low-light videos.By combining motion estimation with reference frame optimization, the initial frame can predict the next 29 frames, reducing inference time by over 80% and computational complexity by 86%. Tests show LLEN accurately extracts edges even with about 80 photons per pixel, enabling EP-SAM to produce precise masks and significantly outperform SAM and SAM-2. EP-SAM improves mean Intersection over Union (mIoU) by 5.85% over SAM on the CamVid dataset. Video demos: https://github.com/wzt22thu/EP-SAM/releases/tag/DEMO. Zhitao Wang, Jiangtao Wen, Yuxing Han 0001 |
ICASSP | 3 |
| 2025 | Smart Grid Powered Datacenters for Computation Offloading: Delay-Sensitive Scheduling and Distributed ProtocolsabstractSmart grids are expected to play a vital role in the era of artificial intelligence (AI), because datacenters are now experiencing a dramatic increase in power consumption, especially with the vast deployments and applications of large language models (LLM). As a result, efficient and timely offloading of computational tasks such as model training in smart grid powered datacenters becomes a challenging but yet critical issue that has attracted considerable recent attention. In particular, our aim is to minimize the overall cost given the time-varying power supply along with its price induced by the dynamic nature of renewable energy, while satisfying a delay constraint that assures the quality-of-service (QoS). To achieve this goal, we present a mixed integer programming (MIP) for cost-efficient and latency-sensitive computation offloading in smart grid powered datacenters. A distributed algorithm for solving the MIP is judiciously conceived, which allows the datacenters and data owners to efficiently schedule the transmission of raw data and the computation for model training in a decentralized manner. Simulations demonstrate that the distributed protocol is capable of adapting to the dynamic energy price, time-varying supply of renewable energy, and even the line outage in smart grids with negligible latency and small signaling overhead. Wei Chen 0002, Yuxing Han 0001 |
ICC | 3 |
| 2025 | ArchiSet: Benchmarking Editable and Consistent Single-View 3D Reconstruction of Buildings with Specific Window-to-Wall Ratios
Pengyu Zeng, Licheng Shen, Miao Zhang 0010, Yuxing Han 0001 |
ICCV | 6 |
| 2025 | A Zero Decoding Approach to Video ClassificationabstractClassifying videos into distinct categories, such as Sport and Music Video, is crucial for multimedia understanding and retrieval, especially with growing content volume. Traditional methods require video decompression to extract pixel-level features like color, texture, and motion, thereby increasing computational and storage demands. We present a novel approach that examines only the compressed bitstream of a video to perform classification, eliminating the need for bitstream decoding. To validate our approach, we built a comprehensive data set comprising over 29,000 YouTube video clips, totaling 6,000 hours and spanning 11 distinct categories. Our evaluations indicate precision, accuracy, and recall rates consistently above 80%, many exceeding 90%, and some reaching 99%. The algorithm operates approximately 15,000 times faster than real-time for 30fps videos, outperforming traditional Dynamic Time Warping (DTW) algorithm by seven orders of magnitude and state-of-the-art video classification model by three orders of magnitude. Chen Ye Gan, Jiangtao Wen, Yuxing Han 0001 |
ICME | 3 |
| 2025 | BackSlash: Rate Constrained Optimized Training of Large Language ModelsabstractThe rapid advancement of large-language models (LLMs) has driven extensive research into parameter compression after training has been completed, yet compression during the training phase remains largely unexplored. In this work, we introduce Rate-Constrained Training (BackSlash), a novel training-time compression approach based on rate-distortion optimization (RDO). BackSlash enables a flexible trade-off between model accuracy and complexity, significantly reducing parameter redundancy while preserving performance. Experiments in various architectures and tasks demonstrate that BackSlash can reduce memory usage by 60\% - 90\% without accuracy loss and provides significant compression gain compared to compression after training. Moreover, BackSlash proves to be highly versatile: it enhances generalization with small Lagrange multipliers, improves model robustness to pruning (maintaining accuracy even at 80\% pruning rates), and enables network simplification for accelerated inference on edge devices. Jiangtao Wen, Yuxing Han 0001 |
ICML | 3 |
| 2025 | A Theoretical Analysis of Efficiency Constrained Utility-Privacy Bi-Objective Optimization in Federated LearningabstractFederated learning (FL) enables multiple clients to collaboratively learn a shared model without sharing their individual data. Concerns about utility, privacy, and training efficiency in FL have garnered significant research attention. Differential privacy has emerged as a prevalent technique in FL, safeguarding the privacy of individual user data while impacting utility and training efficiency. Within Differential Privacy Federated Learning (DPFL), previous studies have primarily focused on the utility-privacy trade-off, neglecting training efficiency, which is crucial for timely completion. Moreover, differential privacy achieves privacy by introducing controlled randomness (noise) on selected clients in each communication round. Previous work has mainly examined the impact of noise level ($\sigma$) and communication rounds ($T$) on the privacy-utility dynamic, overlooking other influential factors like the sample ratio ($q$, the proportion of selected clients). This paper systematically formulates an efficiency-constrained utility-privacy bi-objective optimization problem in DPFL, focusing on$\sigma$,$T$, and$q$. We provide a comprehensive theoretical analysis, yielding analytical solutions for the Pareto front. Extensive empirical experiments verify the validity and efficacy of our analysis, offering valuable guidance for low-cost parameter design in DPFL. Hanlin Gu, Gongxi Zhu, Yuxing Han 0001, Yan Kang 0001, Lixin Fan, Qiang Yang 0001 |
IEEE Trans. Big Data | 4 |
| 2025 | A Hybrid Self-Supervised Learning Framework for Vertical Federated LearningabstractVertical federated learning (VFL), a variant of Federated Learning (FL), has recently drawn increasing attention as the VFL matches the enterprises' demands of leveraging more valuable features to achieve better model performance. However, conventional VFL methods may run into data deficiency as they exploit only aligned and labeled samples (belonging to different parties), leaving often the majority of unaligned and unlabeled samples unused. The data deficiency hampers the effort of the federation. In this work, we propose a Federated Hybrid Self-Supervised Learning framework, named FedHSSL, that utilizes cross-party views (i.e., dispersed features) of samples aligned among parties and local views (i.e., data augmentation) of unaligned samples within each party to improve the representation learning capability of the VFL joint model. FedHSSL further exploits invariant features across parties to boost the performance of the joint model through partial model aggregation. FedHSSL, as a framework, can work with various representative SSL methods. We empirically demonstrate that FedHSSL methods outperform baselines by large margins. We provide an in-depth analysis of FedHSSL regarding label leakage, which is rarely investigated in existing self-supervised VFL works. The experimental results show that, with proper protection, FedHSSL achieves the best privacy-utility trade-off against the state-of-the-art label inference attack compared with baselines. Code is available athttps://github.com/jorghyq2016/FedHSSL. Yuanqin He, Yan Kang 0001, Jiahuan Luo, Lixin Fan, Yuxing Han 0001, Qiang Yang 0001 |
IEEE Trans. Big Data | 6 |
| 2024 | Judging a video by its bitstream coverabstractClassifying videos into distinct categories, such as Sport and Music Video, is crucial for multimedia understanding and retrieval. Traditional methods require video decompression to extract pixel-level features like color, texture, and motion, thereby increasing computational and storage demands. We introduce a novel direction for video classification that does not rely on pixel domain information. Instead, we use the sequence of video frame sizes extracted from compressed bitstreams as input for a ResNet-based deep neural network, without the need for bitstream decoding or parsing. This approach leverages information captured by modern video compression algorithms, particularly advanced spatial and temporal prediction methods found in modern video coding standards such as H.264/AVC, H.265.HEVC and H.266/VVC. Yuxing Han 0001, Yunan Ding, Chen Ye Gan, Jiangtao Wen |
DCC | 1 |
| 2024 | Joint Quickest Line Outage Detection and Emergency Demand Response for Smart GridsabstractEmergency Demand Response Programs (EDRP) have garnered significant attention for its potential to enhance the safety and reliability of smart grids. However, its effectiveness is often compromised by latency in detecting line outages due to noise in observed sources, limited sensor sampling rates, and constrained communication bandwidth. While advancements in quickest line outage detection have been made, the demand reduction actions initiated only after a detected change point may be too late as the EDRP is also time consuming, thereby potentially resulting in severe failures in smart grids. To address these challenges and attain high-assurance power systems, we present a joint framework that integrates quickest line outage detection with emergency demand response, allowing for simultaneous or parallel change point detection and demand reduction. In particular, a Constrained Markov Decision Process (CMDP) is formulated to maximize the average expected reward that characterizes both the revenue and risk of the smart grid, considering operational constraints related to its demand response capabilities. Both theoretical and numerical results demonstrate that our proposed integrated detection and control architecture outperforms the conventional layered successive implementations, achieving superior profit maximization and risk mitigation. Zuxu Chen, Wei Chen 0002, Changkun Li, Yuxing Han 0001 |
GLOBECOM | 4 |
| 2024 | Improving Language Model-Based Zero-Shot Text-to-Speech Synthesis with Multi-Scale Acoustic PromptsabstractZero-shot text-to-speech (TTS) synthesis aims to clone any unseen speaker’s voice without adaptation parameters. By quantizing speech waveform into discrete acoustic tokens and modeling these tokens with the language model, recent language model-based TTS models show zero-shot speaker adaptation capabilities with only a 3-second acoustic prompt of an unseen speaker. However, they are limited by the length of the acoustic prompt, which makes it difficult to clone personal speaking style. In this paper, we propose a novel zero-shot TTS model with the multi-scale acoustic prompts based on a language model. A speaker-aware text encoder is proposed to learn the personal speaking style at the phoneme-level from the style prompt consisting of multiple sentences. Following that, a VALL-E based acoustic decoder is utilized to model the timbre from the timbre prompt at the frame-level and generate speech. The experimental results show that our proposed method outperforms baselines in terms of naturalness and speaker similarity, and can achieve better performance by scaling out to a longer style prompt1. Shun Lei, Yixuan Zhou 0002, Liyang Chen, Zhiyong Wu 0001, Xixin Wu, Shiyin Kang, Yahui Zhou, Yuxing Han 0001, Helen M. Meng |
ICASSP | 10 |
| 2024 | Unlearning during Learning: An Efficient Federated Machine Unlearning Method
Hanlin Gu, Gongxi Zhu, Jie Zhang 0078, Yuxing Han 0001, Lixin Fan, Qiang Yang 0001 |
IJCAI | 5 |
| 2024 | Spatiotemporal smoothing aggregation enhanced multi-scale residual deep graph convolutional networks for skeleton-based gait recognition
Guanghai Chen, Chengzhi Zheng, Junshu Wang, Xinchao Liu, Yuxing Han 0001 |
Appl. Intell. | 6 |
| 2024 | Optimizing Privacy, Utility, and Efficiency in a Constrained Multi-Objective Federated Learning FrameworkabstractConventionally, federated learning aims to optimize a single objective, typically the utility. However, for a federated learning system to be trustworthy, it needs to simultaneously satisfy multiple objectives, such as maximizing model performance, minimizing privacy leakage and training costs, and being robust to malicious attacks. Multi-Objective Optimization (MOO) aiming to optimize multiple conflicting objectives simultaneously is quite suitable for solving the optimization problem of Trustworthy Federated Learning (TFL). In this article, we unify MOO and TFL by formulating the problem of constrained multi-objective federated learning (CMOFL). Under this formulation, existing MOO algorithms can be adapted to TFL straightforwardly. Different from existing CMOFL algorithms focusing on utility, efficiency, fairness, and robustness, we consider optimizing privacy leakage along with utility loss and training cost, the three primary objectives of a TFL system. We develop two improved CMOFL algorithms based on NSGA-II and PSL, respectively, to effectively and efficiently find Pareto optimal solutions and provide theoretical analysis on their convergence. We design quantitative measurements of privacy leakage, utility loss, and training cost for three privacy protection mechanisms: Randomization, BatchCrypt (an efficient homomorphic encryption), and Sparsification. Empirical experiments conducted under the three protection mechanisms demonstrate the effectiveness of our proposed algorithms. Yan Kang 0001, Hanlin Gu, Xingxing Tang, Yuanqin He, Yuzhu Zhang, Jinnan He, Yuxing Han 0001, Lixin Fan, Kai Chen 0005, Qiang Yang 0001 |
ACM Trans. Intell. Syst. Technol. | 7 |
| 2023 | DarkFeat: Noise-Robust Feature Detector and Descriptor for Extremely Low-Light RAW ImagesabstractLow-light visual perception, such as SLAM or SfM at night, has received increasing attention, in which keypoint detection and local feature description play an important role. Both handcraft designs and machine learning methods have been widely studied for local feature detection and description, however, the performance of existing methods degrades in the extreme low-light scenarios in a certain degree, due to the low signal-to-noise ratio in images. To address this challenge, images in RAW format that retain more raw sensing information have been considered in recent works with a denoise-then-detect scheme. However, existing denoising methods are still insufficient for RAW images and heavily time-consuming, which limits the practical applications of such scheme. In this paper, we propose DarkFeat, a deep learning model which directly detects and describes local features from extreme low-light RAW images in an end-to-end manner. A novel noise robustness map and selective suppression constraints are proposed to effectively mitigate the influence of noise and extract more reliable keypoints. Furthermore, a customized pipeline of synthesizing dataset containing low-light RAW image matching pairs is proposed to extend end-to-end training. Experimental results show that DarkFeat achieves state-of-the-art performance on both indoor and outdoor parts of the challenging MID benchmark, outperforms the denoise-then-detect methods and significantly reduces computational costs up to 70%. Code is available at https://github.com/THU-LYJ-Lab/DarkFeat. Yubin Hu 0001, Wang Zhao 0001, Jisheng Li, Yong-Jin Liu 0001, Yuxing Han 0001, Jiangtao Wen |
AAAI | 6 |
| 2023 | Efficient Semantic Segmentation by Altering Resolutions for Compressed VideosabstractVideo semantic segmentation (VSS) is a computationally expensive task due to the per-frame prediction for videos of high frame rates. In recent work, compact models or adaptive network strategies have been proposed for efficient VSS. However, they did not consider a crucial factor that affects the computational cost from the input side: the input resolution. In this paper, we propose an altering resolution framework called AR-Seg for compressed videos to achieve efficient VSS. AR-Seg aims to reduce the computational cost by using low resolution for non-keyframes. To prevent the performance degradation caused by downsampling, we design a Cross Resolution Feature Fusion (CR-eFF) module, and supervise it with a novel Feature Similarity Training (FST) strategy. Specifically, CReFF first makes use of motion vectors stored in a compressed video to warp features from high-resolution keyframes to low-resolution non-keyframes for better spatial alignment, and then selectively aggregates the warped features with local attention mechanism. Furthermore, the proposed FST supervises the aggregated features with high-resolution features through an explicit similarity loss and an implicit constraint from the shared decoding layer. Extensive experiments on CamVid and Cityscapes show that AR-Seg achieves state-of-the-art performance and is compatible with different segmentation backbones. On CamVid, AR-Seg saves 67% computational cost (measured in GFLOPs) with the PSPNet18 back-bone while maintaining high segmentation accuracy. Code: https://github.com/THU-LYJ-Lab/AR-Seg. Yubin Hu 0001, Yanghao Li, Jisheng Li, Yuxing Han 0001, Jiangtao Wen, Yong-Jin Liu 0001 |
CVPR | 5 |
| 2023 | Path Planning Optimization With Multiple Pesticide and Power Loading Bases Using Several Unmanned Aerial Systems on Segmented Agricultural FieldsabstractWe propose a hybrid algorithm based on the nested genetic algorithm (GA) and integer particle swarm optimization (PSO) for multiple unmanned aerial systems (multi-UASs) plant protection operations in multiple segmented fields, with multibases loading pesticides and power. In the proposed algorithm, a preprocessing phase is adopted that includes segmented fields splitting and merging, to shrink the operation field number (i.e., multi-UASs total sorties) and ensure that each generated field area remains smaller than (but close to) the unmanned aerial system (UAS) maximum single-sortie operation area. An improved one-sortie coverage path generation model is established, avoiding power waste in the case when UAS needs to operate on more than one subfield in one sortie. The integer PSO is then used to optimize the initial assignment for UASs on multibases, with the purpose of reducing total operation time and nonoperation flight distance of overall coverage path planning. The proposed algorithm solves the problems of how to design pesticide spraying coverage path on segmented agricultural fields with several bases, and allocate to multiple plant protection UASs with constrained loading capacity. Computer simulations are provided to validate the effectiveness of our algorithms and the performance compared with other algorithms. The test results show that the total number of flight sorties using nested GA (53) is smaller than that using some representative metaheuristic algorithms (57), density-based spatial clustering of applications with noise (60), and traditional planning methods (83). Both nonoperation distance and total operation time obtained by nest GA and integer PSO are smaller than those using other planning methods. The appearance of multibases loading pesticide and power, would reduce the nonspraying flight energy and operation time wastage of UASs. Yang Xu 0030, Yuxing Han 0001, Zhu Sun 0003, Yongkui Jin, Xinyu Xue, Yubin Lan |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2022 | Lyapunov Drift-Based Scheduling for Short-Packet Transmission with Finite Blocklength CodingabstractUltra-Reliable and Low-Latency Communication (URLLC) has attracted considerable attention because it has great potential in industrial automation and autonomous driving. As one of the promising techniques to solve the low-latency problem, finite blocklength coding has been a critical topic. How to reduce the delay with the finite blocklength coding in the URLLC system is a challenging issue. In this paper, a Lyapunov drift-based scheduling scheme is presented for short-packet transmission systems. Specifically, a scheduling scheme is designed to determine the transmission rate and maximize the system throughput with Lyapunov optimization. The average power constraint can be tackled with the virtual power queue method. So that the short-packet transmission problem can be formulated as a univariate optimization problem, which is non-convex. A sub-optimal solution to this optimization problem can be obtained by our presented algorithm, which ignores the higher order terms of the Taylor series. Through simulations, the performance of our proposed algorithm is shown to be very close to that of the optimal policy obtained by the exhaustive search algorithm. Yuanrui Liu, Yuxing Han 0001, Wei Chen 0002 |
GLOBECOM | 3 |
| 2022 | Vision Perception Unit: Next-Generation Smart CMOS Image SensorabstractAs we reach the end of Moore’s Law and Dennard Scaling, it has become highly desirable to design a highly integrated and optimized pipeline specifically for computer vision. A new generation of integrated "smart" visual processors that streamline an end-to-end optimized visual information acquisition and processing pipeline (VIAPP) becomes necessary to lower the cost, power consumption, and latency.We describe a new paradigm for VIAPP as Vision Perception Unit (VPU), wherein electric signals generated by photons are amplified before converting to the digital signals to emulate an initial layer of a convolutional neural network (CNN). The outputs from these layers are then converted to digital signals and processed by following layers of a deep CNN. Wenqi Ji, Yuxing Han 0001, Jiangtao Wen, Yubin Hu 0001, Futang Wang, Jun Zhang 0006 |
HCS | 2 |
| 2022 | Rate Control for Learned Video CompressionabstractRate control is a critical part for video compression, especially in bandwidth-limited tasks such as live and broadcast. The newly-rising learned video compression has shown advantageous rate-distortion (RD) performance in previous research, but lack of rate control heavily limits its usage in real coding scenarios. In this work, we present the first rate control scheme tailored for learned video compression. Specifically, we explore the inter-frame dependency of learned video compression and propose a novel R-D-λ model accordingly for efficient rate allocation. Additionally, a staged update algorithm is developed for robust parameter estimation. Experiments on public datasets show that, the proposed rate control scheme achieves low rate error while maintaining equal or even higher RD performance, without introducing coding time overhead. Yanghao Li, Jisheng Li, Jiangtao Wen, Yuxing Han 0001, Shan Liu 0001, Xiaozhong Xu |
ICASSP | 5 |
| 2021 | Learning Model-Blind Temporal Denoisers without Ground TruthsabstractDenoisers trained with synthetic noises often fail to cope with the diversity of real noises, giving way to methods that can adapt to unknown noise without noise modeling or ground truth. Previous image-based method leads to noise overfitting if directly applied to temporal denoising, and has inadequate temporal information management especially in terms of occlusion and lighting variation. In this paper, we propose a general framework for temporal denoising that successfully addresses these challenges. A novel twin sampler assembles training data by decoupling inputs from targets without altering semantics, which not only solves the noise overfitting problem, but also generates better occlusion masks by checking optical flow consistency. Lighting variation is quantified based on the local similarity of aligned frames. Our method consistently outperforms the prior art by 0.6-3.2dB PSNR on multiple noises, datasets and network architectures. State-of-the-art results on reducing model-blind video noises are achieved. Yanghao Li, Bichuan Guo, Jiangtao Wen, Zhen Xia, Shan Liu 0001, Yuxing Han 0001 |
ICASSP | 6 |
| 2021 | Learning To Compose 6-DOF Omnidirectional Videos Using Multi-Sphere ImagesabstractOmnidirectional video is an essential component of Virtual Reality. Although various methods have been proposed to generate content that can be viewed with six degrees of freedom (6-DoF), existing systems usually involve complex depth estimation, image inpainting or stitching pre-processing. In this paper, we propose a system that uses a 3D ConvNet to generate a multi-sphere images (MSI) representation that can be experienced in 6-DoF VR. The system utilizes conventional omnidirectional VR camera footage directly without the need for a depth map or segmentation mask, thereby significantly simplifying the overall complexity of the 6-DoF omnidirectional video composition. By using a newly designed weighted sphere sweep volume (WSSV) fusing technique, our approach is compatible with most panoramic VR camera setups. A ground truth generation approach for high-quality artifact-free 6-DoF contents is proposed and can be used by the research and development community for 6-DoF content generation. Jisheng Li, Yubin Hu 0001, Yuxing Han 0001, Jiangtao Wen |
ICIP | 4 |
| 2021 | Extending 6-DoF VR Experience Via Multi-Sphere Images InterpolationabstractThree-degrees-of-freedom (3-DoF) omnidirectional imaging has been widely used in various applications ranging from street maps to 3-DoF VR live broadcasting. Although allowing for navigating viewpoints rotationally inside a virtual world, it does not provide motion parallax key for human 3D perception. Recent research mitigates this problem by introducing 3 transitional degrees of freedom (6-DoF) using multi-sphere images (MSI) which is beginning to show promises in handling occlusions and reflective objects. However, the design of MSI naturally limits the range of authentic 6-DoF experiences, as existing mechanisms for MSI rendering cannot fully utilize multi-layer information when synthesizing novel views between multiple MSIs. To tackle this problem and extend the 6-DoF range, we propose an MSI interpolation pipeline that utilizes adjacent MSIs' 3D information embedded inside their layers. In this work, we describe an MSI projection scheme along with an MSI interpolation network to predict intermediate MSIs in order to facilitate the need for extended range. We demonstrate that our system significantly improves the range of 6-DoF experience compared with other MSI-based methods. With extensive experiments, we show our algorithm outperforms state-of-the-art methods both qualitatively and quantitatively in synthesizing novel view panoramas. Jisheng Li, Jinghui Jiao, Yubin Hu 0001, Yuxing Han 0001, Jiangtao Wen |
ACM Multimedia | 5 |
| 2020 | Deep Material Recognition in Light-Fields via Disentanglement of Spatial and Angular Information
Bichuan Guo, Jiangtao Wen, Yuxing Han 0001 |
ECCV (24) | 3 |
| 2020 | Multimodal Video Saliency Analysis With User-Biased InformationabstractVideo saliency is widely used in various video understanding and processing related applications. Despite the fact that studies have indicated the influence of user preferences on visual attention when watching videos, current researches on saliency are based on visual contents and have not taken viewer-related information into account. In this paper, we propose a learning-based multimodal framework to predict video saliency aided by social data analysis. We introduce a popularity assisted attention mechanism into a content-specific neural network to extract spatio-motion features, and utilize a convolutional long short-term memory (ConvLSTM) network to discover temporal characteristics. Experiments demonstrate that our approach outperforms the state-of-the-art video saliency analysis methods, which validates the effectiveness of incorporating external user-biased information into saliency prediction. Jiangyue Xia, Jingqi Tian, Jiangtao Wen, Yuxing Han 0001 |
ICME | 6 |
| 2020 | Genetic Algorithm Based Rate Control for AV1abstractRate control persists as a core problem in video coding area. This paper proposes a robust rate control framework for the recently released AV1 standard, which is a royalty-free video codec specification generated by AOM. Firstly, an exponential rate control model as well as its initialization process and update strategy is proposed to characterize the correspondences between rate and quantization parameter. Next, this paper proposes RC-EMD to measure the similarity between two video frames. Then based on the RC-EMD, a dynamic bit allocation method using genetic algorithm is proposed for AV1 encoder, which can simultaneously detect scene changes. Unlike the existing work, the proposed method explores a larger solution space and automatically adapts to actual input videos, which results in improved performance. Meiyuan Fang, Yuxing Han 0001, Jiangtao Wen |
IEEE Signal Process. Lett. | 2 |
| 2019 | A Conditional Bayesian Block Structure Inference Model for Optimized AV1 EncodingabstractAV1, a next-generation open-source and royalty-free video coding standard, achieves high compression performance at high computational cost. To meet the requirements of HD and UHD video applications, extensive optimizations in both the algorithm and implementation of AV1 are required. In this paper, we analyze the similarities between the block structure decisions after rate-distortion (RD) optimized AV1 and HEVC encodings of the same input. Taking advantage of such similarities, we propose a conditional Bayesian inference model to perform early termination in block partition determination of AV1 based on HEVC encoding outputs. An estimation algorithm is designed to iteratively calculate the prior probability for Bayesian inference. Experiment results show that our proposed algorithm could realize an average time saving of 35.7% and negligible BD-rate loss (0.61%), with the pre-encoding time taken into consideration. Bichuan Guo, Minhao Tang, Yuxing Han 0001, Jiangtao Wen |
ICME | 4 |
| 2019 | AGEM: Solving Linear Inverse Problems via Deep Priors and SamplingabstractIn this paper we propose to use a denoising autoencoder (DAE) prior to simultaneously solve a linear inverse problem and estimate its noise parameter. Existing DAE-based methods estimate the noise parameter empirically or treat it as a tunable hyper-parameter. We instead propose autoencoder guided EM, a probabilistically sound framework that performs Bayesian inference with intractable deep priors. We show that efficient posterior sampling from the DAE can be achieved via Metropolis-Hastings, which allows the Monte Carlo EM algorithm to be used. We demonstrate competitive results for signal denoising, image deblurring and image devignetting. Our method is an example of combining the representation power of deep learning with uncertainty quantification from Bayesian statistics. Bichuan Guo, Yuxing Han 0001, Jiangtao Wen |
NeurIPS | 2 |
| 2019 | SMER: a secure method of exchanging resources in heterogeneous internet of things
Yuxing Han 0001, Jiangtao Wen |
Frontiers Comput. Sci. | 2 |
| 2019 | Hadamard Transform-Based Optimized HEVC Video CodingabstractThe High Efficiency Video Coding (HEVC/H.265) standard achieves great improvement in compression efficiency over the widely used H.264/AVC standard at a cost of much higher complexity. When encoding videos using HEVC, the selection of the quantization parameter (QP) can significantly affect the coding efficiency. Typical algorithms for adaptive quantization employ fixed bitrate budgeting or fixed QP offsets for different frames and different blocks, without considering detailed input video characteristics, some of which might be captured using computer vision methods. The problem of determining adaptive settings of HEVC coding parameters has not been satisfactorily solved. In this paper, we proposed a Hadamard (“HAD”) energy-based optimized HEVC video encoder, in which HAD energy is used to measure the amount of residual information to be encoded in a block so as to determine the QP value for each block for a better coding efficiency. HAD energy is also used to expedite the time-consuming mode decision process and for a precise scene change detection to avoid flicker artifacts. Experiment using the widely used open source HEVC encoder x265-v1.8 showed that the proposed algorithm was able to achieve an average of 14% saving in Bjøntegaard-delta-rate and an average of 20% saving in encoding time as compared the “medium” preset of x265, while the proposed algorithm also produced an improvement of 3.3% in coding efficiency for the HEVC reference software HM-16.6. Minhao Tang, Jiangtao Wen, Yuxing Han 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2019 | A Universal Optical Flow Based Real-Time Low-Latency Omnidirectional Stereo Video SystemabstractOmnidirectional stereoscopic video (ODSV) is a key element of creating an immersive experience for virtual reality that has attracted extensive interest while presenting many technical challenges. Two such key challenges are real-time, low-latency high-quality seamless video stitching from multiple cameras, and faithful reconstruction of 3-D information. Even though various attempts have been made to achieve different combinations of real-time, low-latency, automation, and high output resolution in stereoscopic panoramic video communication, achieving these characteristics simultaneously remains a challenge to be tackled. In this paper, we present a universally applicable and practical end-to-end system based on a novel real-time optical flow algorithm to produce high-quality real-time ODSV with reconstructed 3-D depth information at low latency. Through a configurable process, various camera systems can be calibrated and seamlessly stitched together using the proposed system. The stitched 3-D panoramic video is encoded with a standard compliant video encoder that is optimized for panoramic video. Thanks to various optimizations introduced in this paper, the proposed system is capable of producing real-time ODSV of ultra High definition resolution with a glass to glass latency of 2.2 s using a desktop computer with a single Nvidia graphic card. Experiments show that the proposed system achieves an encoding performance superior to existing open-source HEVC implementations and an optical flow estimation performance better than the Facebook algorithm while running two orders of magnitudes faster. Minhao Tang, Jiangtao Wen, Jiawen Gu, Philip Junker, Bichuan Guo, Guansyun Jhao, Yuxing Han 0001 |
IEEE Trans. Multim. | 9 |
| 2018 | A Bayesian Approach to Block Structure Inference in AV1-Based Multi-Rate Video EncodingabstractDue to differences in frame structure, existing multi-rate video encoding algorithms cannot be directly adapted to encoders utilizing special reference frames such as AV1 without introducing substantial rate-distortion loss. To tackle this problem, we propose a novel bayesian block structure inference model inspired by a modification to an HEVC-based algorithm. It estimates the posterior probabilistic distributions of block partitioning, and adapts early terminations in the RDO procedure accordingly. Experimental results show that the proposed method provides flexibility for controlling the tradeoff between speed and coding efficiency, and can achieve an average time saving of 36.1% (up to 50.6%) with negligible bitrate cost. Bichuan Guo, Jiawen Gu, Yuxing Han 0001, Jiangtao Wen |
DCC | 4 |
| 2018 | Convex Optimization Based Bit Allocation for Light Field Compression Under Weighting and Consistency ConstraintsabstractCompared with conventional image and video, light field images introduce the weight channel, as well as the visual consistency of rendered view, information that has to be taken into account when compressing the pseudo-temporal-sequence (PTS) created from light field images. In this paper, we propose a novel frame level bit allocation framework for PTS coding. A joint model that measures weighted distortion and visual consistency, combined with an iterative encoding system, yields the optimal bit allocation for each frame by solving a convex optimization problem. Experimental results show that the proposed framework is effective in producing desired distortion distribution based on weights, and achieves up to 24.7% BD-rate reduction comparing to the default rate control algorithm. Bichuan Guo, Yuxing Han 0001, Jiangtao Wen |
DCC | 2 |
| 2018 | Multi-Representations Encoding Framework for Adaptive Http StreamingabstractAdaptive HTTP streaming requires a video to be encoded at multiple representations of different target bitrates. To achieve both smooth streaming and good quality, the representations need to be encoded with accurate rate control and the best quality possible for the target bitrates. However, in practical applications, accurate rate control and good video quality are very difficult to achieve at the same time with one-pass and real-time encoding required by low latency applications. In this paper, we proposed a multi-representation encoding framework that reused the encoding information from low-quality representations to accelerate and optimize higher bitrate encodings. The proposed framework is implemented on two different rate control models, namely R-λ model in HM-16.3 and complexity model in x264, to demonstrate the universality. To be best of our knowledge, the optimization in compression performance of multi - representation is first proposed in this paper. The proposed algorithm can achieve better video quality with only small latency in video coding. Results show that up to 49.6% BDRate savings and 4.43dB BDPSNR improvement are achieved in HM as compared with independent one-pass encodings. Meanwhile, 12.5% time and 14.65% BDRate savings were observed in x264. Jiawen Gu, Jiangtao Wen, Bichuan Guo, Yuxing Han 0001 |
ICIP | 4 |
| 2018 | Fast Block Structure Determination in Av1-Based Multiple Resolutions Video EncodingabstractThe widely used adaptive HTTP streaming requires an efficient algorithm to encode the same video to different resolutions. In this paper, we propose a fast block structure determination algorithm based on the AV1 codec that accelerates high resolution encoding, which is the bottle-neck of multiple resolutions encoding. The block structure similarity across resolutions is modeled by the fineness of frame detail and scale of object motions, this enables us to accelerate high resolution encoding based on low resolution encoding results. The average depth of a block's co-located neighborhood is used to decide early termination in the RDO process. Encoding results show that our proposed algorithm reduces encoding time by 30.1%-36.8%, while keeping BD-rate low at 0.71%-1.04%. Comparing to the state-of-the-art, our method halves performance loss without sacrificing time savings. Bichuan Guo, Yuxing Han 0001, Jiangtao Wen |
ICME | 2 |
| 2018 | A Deep Convolutional Network Based Supervised Coarse-to-Fine Algorithm for Optical Flow MeasurementabstractThe measurement of optical flow is an important problem in image processing. There are a number of methods available for optical flow estimation, including traditional variational methods, deep learning based supervised/unsupervised methods. In this work, we propose a deep convolutional network (CNN) based supervised coarse-to-fine approach, which is trained in end-to-end fashion. The proposed method is tested on standard optical flow benchmark datasets including Flying Chairs, MPI Sintel Clean and Final, KITTI. Experimental results show that the proposed framework is able to achieve comparable results to previous approaches with much smaller network architecture. Meiyuan Fang, Yanghao Li, Yuxing Han 0001, Jiangtao Wen |
MMSP | 3 |
| 2018 | Adaptive Intra Candidate Selection With Early Depth Decision for Fast Intra Prediction in HEVCabstractTo better exploit spatial correlations in a video frame, the High Efficiency Video Coding (HEVC) standard has adopted a great many more intra prediction modes than H.264/AVC. As a result, the complexity of rate-distortion-optimized (RDO) HEVC intra mode selection is very high. Many techniques have been proposed to expedite the intra mode selection process to achieve a better overall tradeoff between complexity and RD performance. In this paper, two novel techniques for adaptive intra mode candidate selection and bidirectional depth search algorithms are utilized to accelerate intra prediction. The proposed techniques show an average of 63% (up to 67%) time saving with only 1% BD-Rate increase, outperforming most of existing intra prediction algorithms. Jiawen Gu, Minhao Tang, Jiangtao Wen, Yuxing Han 0001 |
IEEE Signal Process. Lett. | 4 |
| 2018 | Accelerating HEVC Encoding Using Early-SplitabstractThe increase in coding efficiency and complexity of high efficiency video coding (HEVC) over H.264 is due to, among other factors, the time needed to find the optimal partition structure among the more flexible encoding modes for the coding units (CUs) and prediction units (PUs). Although many classification-based algorithms have been proposed to expedite the partition decision, the features that can be acquired from current HEVC encoding order are not sufficient to minimize the loss in coding efficiency. In this letter, we proposed an early-split (ES) order for HEVC CU-level encoding, where the encoder checks the split mode before the nonsquare PU partition modes and utilizes the encoding output of the subCUs to expedite subsequent encoding. Experiments show that the proposed algorithm can save 48% of encoding time on average with only about 0.8% loss in coding performance. Minhao Tang, Jiawen Gu, Yuxing Han 0001, Jiangtao Wen, Shiqiang Yang |
IEEE Signal Process. Lett. | 4 |
| 2017 | Early-Split Based Fast HEVC EncodingabstractThe High Efficiency Video Coding (HEVC) standard achieves 50% improvement incompression efficiency over the widely used H.264/AVC standard at a cost of much higher complexity. The increase in complexity is due to, among other factors, the time needed to findthe optimal partition structure among the more flexible possibilities for the coding units (CUs) and prediction units (PUs). Many classification based algorithms have been proposed to reduce this partition decision time, but the features that can be acquired from current HEVC encoding order may not be sufficient to control the loss in coding efficiency. In this paper, we proposed an Early-Split (ES) order for HEVC encoding, where the encoder checks the split mode before the non-square PU partition modes and utilizes the encoding output of the subCUs to expedite subsequent encoding. Experiments show that the proposed algorithm achieved an average of 48% saving in encoding time with only 0.92% loss in the coding performance. Minhao Tang, Jiawen Gu, Yuxing Han 0001, Jiangtao Wen |
DCC | 4 |
| 2017 | HEVC-based motion compensated joint temporal-spatial video denoisingabstractA novel HEVC-based efficient video denoising algorithm is proposed in this paper. It uses a spatial Gaussian filter for the chrominance components and then utilizes the HEVC motion estimation process to find the best temporal correspondence for low-pass filtering. Other HEVC tools such as quantization, the interpolation and the in-loop filters are also used. Experiments implementing the proposed algorithm in the open-source HEVC encoder ×265 showed a good denoising performance with a much lower computing complexity than the competitors. The performance was comparable to those highly sophisticated algorithms such as the VBM4D, which is 200 times slower. The proposed algorithm can be easily integrated into the real-world video processing systems due to its compatibility with the HEVC standard. Minhao Tang, Yuxing Han 0001, Jiangtao Wen, Shiqiang Yang |
ICASSP | 2 |
| 2017 | Probabilistic graphical model based fast HEVC inter predictionabstractThe High Efficiency Video Coding (HEVC) standard achieves 50% improvement in coding efficiency compared with H.264/ AVC by introducing many more video encoding tools achieving different coding performance and complexity tradeoffs [1]. Various techniques have been proposed to reduce the complexity of HEVC encoding. In this paper, an effective Bayesian network based complexity reduction framework for HEVC encoding is proposed. The proposed framework is able to calculate the conditional probabilities of different status in the encoding process, which can be used to speedup the encoders by neglecting small probability events. Experimental results show that the proposed algorithm can save 52.5% of computational complexity with a loss of 1.41% in compression performance on average. Meiyuan Fang, Jiangtao Wen, Yuxing Han 0001 |
ICIP | 3 |
| 2017 | TCP-ACC: performance and analysis of an active congestion control algorithm for heterogeneous networks
Jun Zhang 0006, Jiangtao Wen, Yuxing Han 0001 |
Frontiers Comput. Sci. | 3 |
| 2016 | A novel low delay in-loop filtering WPP process for parallel HEVC encodingabstractWavefront parallel processing (WPP) is a parallelization technique that enables processing of several rows of Largest Coding Units (LCUs) in parallel. It achieves a relatively high level of parallelism with reasonable loss in compression performance. At the same time, the HEVC standard specifies two in-loop filters, namely the deblocking filter and the sample adaptive offset (SAO), to improve subjective quality as well as coding efficiency. Because SAO parameters cannot be precisely determined until the lower right deblocked samples are available, a delay of several rows is introduced to implement the filtering process. Since the reconstruction of the LCUs will not start until finish filtering, the total delay of WPP and filtering is at least two rows, which obviously influences the performance of parallel processing. In this paper, we propose a novel In-Loop Filtering WPP method that reduces the row delay into four LCUs and significantly improves the parallelism with little rate-distortion (RD) performance loss. Experimental results show that the proposed algorithms can achieve up to the 2.89× speedup compared with the existing WPP method with 24-core server, where the speedup improves with the increase of core number. Jiawen Gu, Yuxing Han 0001, Jiangtao Wen |
VCIP | 2 |
| 2016 | TCP-FIT: An improved TCP algorithm for heterogeneous networks
Jingyuan Wang 0001, Jiangtao Wen, Jun Zhang 0006, Zhang Xiong 0001, Yuxing Han 0001 |
J. Netw. Comput. Appl. | 5 |
| 2015 | DC-Vegas: A delay-based TCP congestion control algorithm for datacenter applications
Jingyuan Wang 0001, Jiangtao Wen, Chao Li 0001, Zhang Xiong 0001, Yuxing Han 0001 |
J. Netw. Comput. Appl. | 5 |
| 2014 | Achieving high throughput and TCP Reno fairness in delay-based TCP over large networks
Jingyuan Wang 0001, Jiangtao Wen, Yuxing Han 0001, Jun Zhang 0006, Chao Li 0001, Zhang Xiong 0001 |
Frontiers Comput. Sci. | 3 |
| 2012 | Highly Scalable Parallel Arithmetic Coding on Multi-Core Processors Using LDPC CodesabstractWe describe a highly scalable parallel arithmetic coder for Markov inputs suitable for implementation on modern multi-core processors. The algorithm divides the input into interleaved sub-sequences which can be then processed independently on different processing units using LDPC-based Slepian-Wolf coding. Experimental simulations show good scalability of the proposed algorithm while also maintaining good compression performance. Notably, when compared with traditional parallel arithmetic coding, the proposed method maintains a much higher efficiency both respect to the entropy limit as well as in terms of the ability to distribute computations across multiple cores without performance loss. WeiDong Hu, Jiangtao Wen, Weiyi Wu, Yuxing Han 0001, Shiqiang Yang, John D. Villasenor |
IEEE Trans. Commun. | 4 |
| 2012 | Efficient Video Coding Using Legacy Algorithmic ApproachesabstractWe show that for high bit rates, a video coding algorithm using a suitable combination of the QM coder and on other methods first published over 20 years ago can deliver video quality rivaling that of H.264 at lower complexity. This has implications both technically, since encoders built using these methods can be more power efficient, and commercially, given the complex licensing and intellectual property issues that accompany newer coding methods such as H.264 and MPEG-4. The methods described in this paper are the basis for the recent decision of the MPEG standards group to begin work on what is referred to as the “Type-1 Video Coding” standard, which, in addition to aiming for high coding efficiency, is intended to minimize royalty issues. John D. Villasenor, Yuxing Han 0001, Yaocheng Rong, Cliff Reader, Jiangtao Wen |
IEEE Trans. Multim. | 5 |
| 2011 | A Compressive Sensing Reconstruction Algorithm for Trinary and Binary Sparse Signals Using Pre-mappingabstractIn this paper, we first analyze impact of the distribution of sparse signals on reconstruction quality in compressive sensing through experimental results and heuristic analysis. We suggest that trinary/binary sparse signals are one of the most difficult signals to reconstruct in terms of error bounds. We then show that by incorporating linear or non-linear mapping prior to sensing, significant improvement in the recovery performance can be achieved. Zhuoyuan Chen, Jiangtao Wen, Jianwei Ma 0006, Yuxing Han 0001, John D. Villasenor |
DCC | 5 |
| 2011 | Automatic SoC design flow on many-core processors: a software hardware co-design approach for FPGAsabstractTraditional FPGA-based system-on-chip (SoC) design in general is accomplished via separate software and hardware design flows. With such a separate design methodology, extra development overhead has to be paid to meet the final system's performance, size and power consumption requirements. To overcome this development overhead which usually leads to significant increase of the time-to-market, a unified and efficient SoC design flow is needed. The current work addresses this problem via a SoC design flow which allows automatic building of a complete autonomous system on an FPGA in accordance with the need of a specific application. In the proposed design flow the architecture of the generated hardware is tailored to match the parallelism granularity and communication structure of the application. This, in turn, allows the application developer to meet the system's performance, size and power consumption requirements with a short time-to-market. To prove the applicability of the proposed approach, a monitor for real-time electrocardiographic (ECG) signal analysis and a motion detection algorithm have been implemented. Oleksii Morozov, Yuxing Han 0001, Jürg Gutknecht, Patrick R. Hunziker |
FPGA | 3 |
| 2011 | TCP-FIT: An improved TCP congestion control algorithm and its performanceabstractThe Transport Control Protocol (TCP) has been widely used by wired and wireless Internet applications such as FTP, email and HTTP. Numerous congestion algorithms have been proposed to improve the performance of TCP in various scenarios, especially for high bandwidth-delay product (BDP) and wireless networks. Although different algorithms may achieve different performance improvements under different network conditions, designing a congestion algorithm that performs well across a wide spectrum of network conditions remains a great challenge. In this paper, we propose a novel congestion control algorithm, named TCP-FIT, which could perform gracefully in both wireless and high BDP networks. The algorithm was inspired by parallel TCP, but with the important distinctions that only one TCP connection with one congestion window is established for each TCP session, and that no modifications to other layers (e.g. the application layer) of the end-to-end system need to be made. Extensive experimental results obtained using both network simulators as well as over “live” wired line, WiFi and 3G networks at different geographical locations and at different times of the day are presented. The performance of the algorithm shown in the experiment results is significantly improved as compared to other state-of-the-art algorithms, while maintaining good fairness. Jingyuan Wang 0001, Jiangtao Wen, Jun Zhang 0006, Yuxing Han 0001 |
INFOCOM | 4 |
| 2011 | Probabilistic Estimation of the Number of Frequency-Hopping TransmittersabstractWe present two probabilistic estimation techniques for identifying the most likely number of frequency-hopping transmitters for a range of different scenarios and compare their performances. In the first technique, cumulative estimation, a Gaussian approximation methodology is developed based on a single integrated measurement over the observation time window. In the second technique, time-distributed estimation, a maximum-likelihood formulation is adopted, and time-specific observation data is used. We give specific analytical consideration to the potential that not all transmitters that are present will be always detected, and explore the effects of the probability of misdetection on the overall estimation process. Simulation results confirm that the approaches presented here can lead with high probability to a correct decision regarding the number of transmitters. Yuxing Han 0001, Jiangtao Wen, Danijela Cabric, John D. Villasenor |
IEEE Trans. Wirel. Commun. | 1 |
| 2011 | Weighted Centroid Localization Algorithm: Theoretical Analysis and Distributed ImplementationabstractInformation about primary transmitter location is crucial in enabling several key capabilities in cognitive radio networks, including improved spatio-temporal sensing, intelligent location-aware routing, as well as aiding spectrum policy enforcement. Compared to other proposed non-interactive localization algorithms, the weighted centroid localization (WCL) scheme uses only the received signal strength information, which makes it simple to implement and robust to variations in the propagation environment. In this paper we present the first theoretical framework for WCL performance analysis in terms of its localization error distribution parameterized by node density, node placement, shadowing variance, correlation distance and inaccuracy of sensor node positioning. Using this analysis, we quantify the robustness of WCL to various physical conditions and provide design guidelines, such as node placement and spacing, for the practical deployment of WCL. We also propose a power-efficient method for implementing WCL through a distributed cluster-based algorithm, that achieves comparable accuracy with its centralized counterpart. Jun Wang 0007, Paulo Urriza, Yuxing Han 0001, Danijela Cabric |
IEEE Trans. Wirel. Commun. | 3 |
| 2010 | Image Compression Using the DCT and Noiselets: A New Algorithm and Its Rate Distortion PerformanceabstractWe describe an image coding algorithm combining the DCT and noiselet information. The algorithm first transmits DCT information sufficient to reproduce a "low-quality" version of the image at the decoder. This image is then used both at the decoder and encoder to create a mutually known list of locations of likely significant noiselet coefficients. The coefficient values themselves are then transmitted to the decoder differentially, by subtracting, at the encoder, the low-quality image from the original image, obtaining the noiselet values and subjecting them to quantization and entropy coding. There remain significant opportunities for further work combining CS-inspired information theoretic techniques with the rate-distortion considerations that are critical in practical image communications. Zhuoyuan Chen, Jiangtao Wen, Shiqiang Yang, Yuxing Han 0001, John D. Villasenor |
DCC | 4 |
| 2010 | Reconstruction of Sparse Binary Signals Using Compressive SensingabstractSummary form only given. This paper has described an improved algorithm for reconstructing sparse binary signals using compressive sensing. The algorithm is based on the reweighted lqnorm optimization algorithm, but with the important additional operation of bounding in each round of the interior-point method iteration, and progressive reduction of q. Experimental results confirm that the algorithm performs well both in terms of the ability to recover an input signal as well as in terms of speed. We also found that both the progressive reduction and the bounding are integral to the improvement in performance. Future work includes extending this approach to Gaussian distributed, as opposed to binary inputs. Jiangtao Wen, Zhuoyuan Chen, Shiqiang Yang, Yuxing Han 0001, John D. Villasenor |
DCC | 4 |
| 2010 | A compressive sensing image compression algorithm using quantized DCT and noiselet informationabstractInspired by recent theoretical advances in compressive sensing (CS), we propose a new framework that combines the classical local discrete cosine transform used in image compression algorithms such as JPEG with a global noiselet measure which is solved using second order cone programming (SOCP). Jiangtao Wen, Zhuoyuan Chen, Yuxing Han 0001, John D. Villasenor, Shiqiang Yang |
ICASSP | 3 |
| 2010 | A Probabilistic Approach to Identifying the Number of Transmitters in the Presence of MisdetectionabstractWe present an analytical framework for identifying the most likely number of frequency-hopping transmitters in the presence of potential misdetection. The problem is formulated for the case where there is a single, global misdetection probability characterizing all transmitter/sensor links, and for the case where the misdetection probabilities are permitted to be different across the different pairwise transmitter/sensor links. Simulation results confirm that the approach can lead with high probability to a correct decision regarding the number of interferers. Yuxing Han 0001, Jiangtao Wen, Danijela Cabric, Sateesh Addepalli, John D. Villasenor |
ICC | 1 |
| 2009 | A Probabilistic Approach to Identifying the Number of Frequency Hoppers for Spectrum SensingabstractCharacterizing the number and type of transmitters occupying a given frequency band is a critical aspect of spectrum sensing specifically and cognitive radio generally. We present an analytical framework based on probability to identify the number of frequency hopping transmitters of one specific type in a band of interest, and show that the probability mass functions associated with the different potential number of transmitters quickly becomes Gaussian as the number of channel observations increases. Simulation results confirm that the approach can lead with high probability to a correct decision regarding the number of interferers. Thus, the methods here can serve as a valuable complement to other spectrum sensing approaches. Yuxing Han 0001, Shaunak Joshi, Lillian L. Dai, Danijela Cabric, Sateesh Addepalli, Jiangtao Wen, John D. Villasenor |
GLOBECOM | 1 |