VLDB 2026 Research / reviewers in the wild / expert
Wenjun Zhang 0001
dblp:46/3359-1
· DBLP profile ↗
287ranked-venue papers
2as first author
142since 2021 · last 2026
0000-0001-8799-1182ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 138 · 1 first-author · 51 since 2021Computer networks · 86 · 1 first-author · 69 since 2021Artificial intelligence and machine learning · 50 · 26 since 2021Systems, architecture and hardware · 19 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021Databases, data management, data science and information retrieval · 5 · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FDD CSI Feedback under Finite Downlink Training: A Rate-Distortion Perspective
Shuao Chen, Junyuan Gao, Yuxuan Shi 0001, Yongpeng Wu 0001, Giuseppe Caire, H. Vincent Poor, Wenjun Zhang 0001 |
ICC | 7 |
| 2026 | ACE-Grouped Neural Min-Sum Decoding with a GRU Hypernetwork for LDPC Codes
Yin Xu 0001, Hao Ju 0002, Dazhi He, Wenjun Zhang 0001 |
ICC | 6 |
| 2026 | On the Fundamental Tradeoff of Sensing Accuracy, Outage Capacity and Information Freshness in ISAC Systems
Zijin Wang, Junyuan Gao, Yongpeng Wu 0001, Wenjun Zhang 0001 |
ICC | 4 |
| 2026 | Multi-hop Parallel Image Semantic Communication for Distortion Accumulation MitigationabstractExisting semantic communication schemes primarily focus on single-hop scenarios, overlooking the challenges of multi-hop wireless image transmission. As semantic communication is inherently lossy, distortion accumulates over multiple hops, leading to significant performance degradation. To address this, we propose the multi-hop parallel image semantic communication (MHPSC) framework, which introduces a parallel residual compensation link at each hop against distortion accumulation. To minimize the associated transmission bandwidth overhead, a coarse-to-fine residual compression scheme is designed. A deep learning-based residual compressor first condenses the residuals, followed by the adaptive arithmetic coding (AAC) for further compression. A residual distribution estimation module predicts the prior distribution for the AAC to achieve fine compression performances. This approach ensures robust multi-hop image transmission with only a minor increase in transmission bandwidth. Experimental results confirm that MHPSC outperforms both existing semantic communication and traditional separated coding schemes. Bingyan Xie, Jihong Park, Yongpeng Wu 0001, Wenjun Zhang 0001, Tony Q. S. Quek |
ICC | 4 |
| 2026 | MA-Aided Hierarchical Hybrid Beamforming for Multi-User Wideband Beam Squint MitigationabstractIn wideband near-field arrays, frequency-dependent array responses cause wavefronts at different frequencies to deviate from that at the center frequency, producing beam squint and degrading multi-user performance. True-time-delay (TTD) circuits can realign the frequency dependence but require large delay ranges and intricate calibration, limiting scalability. Another line of work explores one- and two-dimensional array geometries, including linear, circular, and concentric circular, that exhibit distinct broadband behaviors such as different beam-squint sensitivities and focusing characteristics. These observations motivate adapting the array layout to enable wideband-friendly focusing and enhance multi-user performance without TTD networks. We propose a movable antenna (MA) aided architecture based on hierarchical sub-connected hybrid beamforming (HSC-HBF) in which antennas are grouped into tiles and only the tile centers are repositioned, providing slow geometric degrees of freedom that emulate TTD-like broadband focusing while keeping hardware and optimization complexity low. We show that the steering vector is inherently frequency dependent and that reconfiguring tile locations improves broadband focusing. Simulations across wideband near-field scenarios demonstrate robust squint suppression and consistent gains over fixed-layout arrays, achieving up to 5\% higher sum rate, with the maximum improvement exceeding 140\%. Cixiao Zhang, Yin Xu 0001, Xinghao Guo, XiaoWu Ou, Dazhi He, Wenjun Zhang 0001 |
ICC | 6 |
| 2026 | Low-Complexity Soft-Feedback Detector for AFDM SystemsabstractAffine frequency division multiplexing (AFDM), an emerging multi-carrier modulation scheme, has garnered significant attention due to its resilience to Doppler shifts and capability to achieve full diversity in doubly dispersive channels. However, existing data detection algorithms for AFDM systems face a significant trade-off between computational complexity and accuracy. In this paper, a novel low-complexity data detection scheme, termed the soft-feedback detector (SFD), is proposed. Particularly, building upon a maximum ratio combining (MRC) estimator framework, the SFD leverages the a priori symbol distribution to mitigate error propagation during iterative detection. Specifically, soft-decision feedback is incorporated as extrinsic information derived from the log-likelihood ratios of the transmitted symbols. As a result, the proposed detector significantly enhances detection accuracy while maintaining low computational complexity. Simulation results demonstrate that the SFD consistently outperforms benchmark decision-feedback detectors. In particular, compared with the conventional MRC detector, the proposed scheme achieves approximately a 3 dB signal-to-noise ratio (SNR) gain at the bit error rate (BER) of $10^{-3}$. Taohe Chen, Yin Xu 0001, Tianyao Ma, Aimin Tang, Qu Luo, Dazhi He, Wenjun Zhang 0001 |
ISIT | 7 |
| 2026 | WVSC: Wireless Video Semantic Communication with Multi-Frame CompensationabstractExisting wireless video transmission schemes directly conduct video coding in pixel level, while neglecting the inner semantics contained in videos. In this paper, we propose a wireless video semantic communication framework, abbreviated as WVSC, which integrates the idea of semantic communication into wireless video transmission scenarios. WVSC first encodes original video frames as semantic frames and then conducts video coding based on such compact representations, enabling the video coding in semantic level rather than pixel level. Moreover, to further reduce the communication overhead, a reference semantic frame is introduced to substitute motion vectors of each frame in common video coding methods. At the receiver, multi-frame compensation (MFC) is proposed to produce compensated current semantic frame with a multi-frame fusion attention module. With both the reference frame transmission and MFC, the bandwidth efficiency improves with satisfying video transmission performance. Experimental results verify the performance gain of WVSC over other DL-based methods e.g. DVSC about 1 dB and traditional schemes about 2 dB in terms of PSNR. Bingyan Xie, Yongpeng Wu 0001, Yuxuan Shi 0001, Biqian Feng, Wenjun Zhang 0001, Jihong Park, Tony Q. S. Quek |
WCNC | 5 |
| 2026 | Cross-Splitting-Based Information Geometry Approach for Xl-Mimo Uplink Detection
Wenjun Zhang 0001, Anan Lu, Xiqi Gao 0001 |
WCNC | 1 |
| 2026 | SNR Analysis and Channel Estimation for Multi-UAV Near-Field Communications
Tianyu Huo, Jian Xiong 0001, Yiyan Wu 0001, Songjie Yang, Bo Liu 0001, Wenjun Zhang 0001 |
IEEE Internet Things J. | 6 |
| 2026 | ICDM: Interference Cancellation Diffusion Models for Wireless Semantic CommunicationsabstractDiffusion models (DMs) have recently achieved significant success in wireless communications systems due to their denoising capabilities. The broadcast nature of wireless signals makes them susceptible not only to Gaussian noise, but also to unaware interference. This raises the question of whether DMs can effectively mitigate interference in wireless semantic communication systems. In this paper, we model the interference cancellation problem as a maximum a posteriori (MAP) problem over the joint posterior probability of the signal and interference, and theoretically prove that the solution provides excellent estimates for the signal and interference. To solve this problem, we develop an interference cancellation diffusion model (ICDM), which decomposes the joint posterior into independent prior probabilities of the signal and interference, along with the channel transition probability. The log-gradients of these distributions at each time step are learned separately by DMs and accurately estimated through deriving. ICDM further integrates these gradients with advanced numerical iteration method, achieving accurate and rapid interference cancellation. Extensive experiments demonstrate that ICDM significantly reduces the mean square error (MSE) and enhances perceptual quality compared to schemes without ICDM. For example, on the CelebA dataset under the Rayleigh fading channel with a signal-to-noise ratio (SNR) of 20 dB and signal to interference plus noise ratio (SINR) of 0 dB, ICDM reduces the MSE by 4.54 dB and improves the learned perceptual image patch similarity (LPIPS) by 2.47 dB. The code is available at https://github.com/Wireless3C-SJTU/ICDM. Tong Wu 0003, Zhiyong Chen 0002, Dazhi He, Feng Yang 0006, Meixia Tao, Xiaodong Xu 0001, Wenjun Zhang 0001, Ping Zhang 0003 |
IEEE J. Sel. Areas Commun. | 7 |
| 2026 | Pragmatic Communication in Multi-Agent Collaborative PerceptionabstractCollaborative perception allows each agent to enhance its perceptual abilities by exchanging messages with others. It inherently results in a trade-off between perception ability and communication costs. Previous works transmit complete full-frame high-dimensional feature maps among agents, resulting in substantial communication costs. To promote communication efficiency, we propose only transmitting the information needed for the collaborator's downstream task. This pragmatic communication strategy focuses on three key aspects: i) pragmatic message selection, which selects task-critical parts from the complete data, resulting in spatially and temporally sparse feature vectors; ii) pragmatic message representation, which achieves pragmatic approximation of high-dimensional feature vectors with a task-adaptive dictionary, enabling communicating with integer indices; iii) pragmatic collaborator selection, which identifies beneficial collaborators, pruning unnecessary communication links. Following this strategy, we first formulate a mathematical optimization framework for the perception-communication trade-off and then propose PragComm, a multi-agent collaborative perception system with two key components: i) single-agent detection and tracking and ii) pragmatic collaboration. The proposed PragComm promotes pragmatic communication and adapts to a wide range of communication conditions. We evaluate PragComm for both collaborative 3D object detection and tracking tasks in both real-world, V2V4Real, and simulation datasets, OPV2V and V2X-SIM2.0. PragComm consistently outperforms previous methods with more than 32.7 K× lower communication volume on OPV2V. Yue Hu 0011, Xianghe Pang, Xiaoqi Qin, Yonina C. Eldar, Siheng Chen, Ping Zhang 0003, Wenjun Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2026 | HMSR: Hypercomplex-guided mamba for fine-texture coupling in single image super-resolution
Fengqian Sun, Qiqi Kou, Deqiang Cheng 0001, Guangtao Zhai, Wenjun Zhang 0001 |
Pattern Recognit. | 8 |
| 2026 | Algebraic FDOA-only geolocation with three observers
Wenjun Zhang 0001, Xi Li 0020, Le Yang 0001, Fucheng Guo 0001 |
Signal Process. | 1 |
| 2026 | Low-Latency Satellite-to-Device Interference Detection: A Statistical Change Detection Approach
Runnan Liu, Weifeng Zhu, Shu Sun 0001, Meixia Tao, Wenjun Zhang 0001 |
IEEE Trans. Commun. | 5 |
| 2026 | Node-Based Soft-Output Fast Successive Cancellation List Decoding of Polar CodesabstractThe soft-output successive cancellation list (SOSCL) decoder provides a methodology for estimating the aposteriori probability log-likelihood ratios by only leveraging the conventional SCL decoder of polar codes. However, the sequential decoding nature of SCL introduces high decoding latency to SOSCL. In this paper, we incorporate node-based fast decoding into the SO-SCL framework. After addressing the challenge of soft output extraction in special node decoding, we proposed the soft-output fast SCL (SO-FSCL) decoding algorithm, along with its log-domain implementation and hardware-friendly version. The proposed SO-FSCL decoder can be regarded as an addon extension to FSCL decoder, enabling us to autonomously choose whether to output only hard decisions like FSCL or to provide additional soft outputs. Latency and complexity analyses demonstrate that SO-FSCL can significantly reduce, for example, decoding time steps by 81.8% (with unlimited resources), the number of additions by 41.3%, and the number of comparisons by 46.4%. Meanwhile, simulation results indicate that SO-FSCL delivers almost the same soft-output performance as SO-SCL, outperforming other soft-output polar decoders, especially in scenarios involving iterative decoding. Yongpeng Wu 0001, Zhen Gao 0001, Yin Xu 0001, Xiaohu You 0001, Xiqi Gao 0001, Wenjun Zhang 0001 |
IEEE Trans. Commun. | 7 |
| 2026 | Scalable GNN-Based Power Allocation for Rate-Splitting Cell-Free Massive MIMO SystemsabstractCell-free massive multiple-input multiple-output (CF-mMIMO) systems provide enhanced coverage and capacity for next-generation wireless networks. However, CF-mMIMO systems face significant challenges in downlink power allocation (PA) due to imperfect channel state information (CSI), severe multi-user interference (MUI), and high computational complexity. To address these issues, rate-splitting multiple access (RSMA) is adopted as a robust interference management strategy. Accordingly, this paper proposes an unsupervised and scalable graph neural network (GNN) framework for PA in rate-splitting CF-mMIMO (RS-CF-mMIMO) systems, relying exclusively on large-scale fading (LSF) coefficients without instantaneous CSI. To resolve the dimensionality mismatch in dynamic networks, we introduce a slice-based adaptive layer that projects variable-dimension features into a fixed latent space. This mechanism enables a unified model to generalize across diverse topologies without retraining. Within this architecture, the sum spectral efficiency (SE) is maximized under per-AP power constraints, assuming maximum-ratio precoding for common streams and regularized zero-forcing precoding for private streams. We also derive a weighted minimum mean-square error-alternating direction method of multipliers (WMMSE-ADMM) algorithm as a performance upper bound. Extensive simulations verify that the proposed GNN framework achieves near-optimal SE and outperforms unsupervised deep neural networks (DNNs) across diverse system sizes and pilot assignment schemes. Furthermore, the scalable variant maintains robust performance while reducing the trainable parameter count by over 57% relative to DNNs and decreasing inference latency by up to three orders of magnitude compared with WMMSE-ADMM. Ruomeng Wang, Yin Xu 0001, Aimin Tang, XiaoWu Ou, Dazhi He, Lifeng Wang 0002, Wenjun Zhang 0001 |
IEEE Trans. Commun. | 8 |
| 2026 | Wireless Video Semantic Communication With Decoupled Diffusion Multi-Frame CompensationabstractExisting wireless video transmission schemes directly conduct video coding in pixel level, while neglecting the inner semantics contained in videos. In this paper, we propose a wireless video semantic communication framework with decoupled diffusion multi-frame compensation (DDMFC), abbreviated as WVSC-D, which integrates the idea of semantic communication into wireless video transmission scenarios. WVSC-D first encodes original video frames as semantic frames and then conducts video coding based on such compact representations, enabling the video coding in semantic level rather than pixel level. Moreover, to further reduce the communication overhead, a reference semantic frame is introduced to substitute motion vectors of each frame in common video coding methods. At the receiver, DDMFC is proposed to generate compensated current semantic frame by a two-stage conditional diffusion process. With both the reference frame transmission and DDMFC frame compensation, the bandwidth efficiency improves with satisfying video transmission performance. Experimental results verify the performance gain of WVSC-D over other DL-based methods e.g. DVSC about 1.8 dB in terms of PSNR. Bingyan Xie, Yongpeng Wu 0001, Yuxuan Shi 0001, Biqian Feng, Wenjun Zhang 0001, Jihong Park, Tony Q. S. Quek |
IEEE Trans. Commun. | 5 |
| 2026 | A Lightweight Deep and Wide Network for Image-Based Detection of Industrial Waste GasabstractDue to inadequate monitoring, key pollutants (e.g., PM2.5, VOCs, etc) very possibly leak into atmosphere, thus to endanger the long-term and short-term life safety of people that work and live in the environment. Therefore, it is imperative to effectively and efficiently detect the leakage of industrial waste gas, for the purpose of timely lowering the risk of pollution and explosions. To solve such a problem, we in this paper propose a new lightweight deep and wide network (LdwNet) for detecting the leakage of industrial waste gas from an image, which brings about the two main merits: 1) Compensating for the deficiencies of sensor-based detection methods, which can accurately detect the leakage of waste gas and even measure its concentrations but require to seek leakage sources beforehand; 2) Overcoming the shortcomings of image-based detection methods, which leverage DNN-based recognition technologies and usually suffer from low efficacy, low efficiency and high energy consumption during the model training and inference. To specify, the proposed LdwNet is developed by simulating human perception, motivated by the method which detects the leakage of industrial waste gas from surveillance images with the human observation and judgement. First, based on the inspiration that the human eyes are highly sensitive to horizontal and vertical stimuli, we construct a novel lightweight parallel-series-stripe (PS2) module to validly extract features with very few parameters. Second, to fully exploit deep and shallow features for fusing the global and local information, we extend the PS2 module as a backbone along both the deep and wide directions to build the multi-channel network. Third, to achieve effective, efficient and low-carbon detection in model running, we constraint the extended PS2 modules with parameter sharing to prodigiously reduce the model parameters and thus to make the proposed model ultra-lightweight. Experiments on the datasets of carbon particulate matters and ethylene leakage prove that our LdwNet with ten thousand parameters outperforms the state-of-the-art models with millions of parameters in detection accuracy and implementation cost, and this renders our proposed LdwNet more suitable for real industrial applications. Ke Gu 0001, Hongyan Liu 0004, Jingchao Cao, Lai-Kuan Wong, Junfei Qiao 0001, Guangtao Zhai, Wenjun Zhang 0001, Weisi Lin, Sam Kwong |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2026 | Distilling Complexity-Scalable Learned Image Compression Models via Neural Architecture Search
Shen Wang 0013, Zhengxue Cheng, Donghui Feng 0003, Cheems Wang, Qunshan Gu, Li Song 0001, Wenjun Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2026 | Diff-Restorer: Unleashing Visual Prompts for Diffusion-Based Universal Image RestorationabstractImage restoration aims to recover high-quality images from degraded observations, yet real-world degradations are complex, coupled, and difficult to model. Existing task-specific methods struggle to generalize beyond predefined degradation types, while recent all-in-one or prompt-based methods still face three key challenges: (1) they rely on task-specific training or fixed prompt pools, limiting adaptability to real-world and mixed degradations; (2) human-instruction or implicit-prompt mechanisms make them difficult to use in practice; and (3) they often fail to balance structural fidelity and perceptual realism. To address these issues, we propose Diff-Restorer, a diffusion-based universal image restoration framework that unifies diverse degradation handling within a single model. Diff-Restorer adaptively extracts decoupled visual prompts from a visual-language model (CLIP), including clear semantic and degradation embeddings. The clear semantic embeddings serve as content prompts to guide the diffusion model for generation, improving perceptual quality. The degradation embeddings as the task identifier modulate the Image-guided Control Module to generate structure control, ensuring faithfulness. Furthermore, we design a Task-aware Decoder to perform structural correction and convert the latent code to the pixel domain. Extensive experiments on various single, real-world, and mixed degradation tasks show that Diff-Restorer outperforms state-of-the-art methods in terms of generality, realism, and fidelity. Hengsheng Zhang, Xinning Chai, Zhengxue Cheng, Rong Xie 0004, Li Song 0001, Wenjun Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2026 | Infrared Image Quality Estimation With Node-to-Graph RegressionabstractBy comparison with the commonly seen visible light images that can be effectively characterized within a Euclidean space, infrared images have non-Euclidean characteristics since their pixels contain rich thermal radiation information, such as heat distribution, surface temperature and thermal radiation. Considering the advantages of Graph Convolutional Networks (GCNs) in processing non-Euclidean data, this study proposes to introduce the GCNs to estimate the quality of infrared images by developing the Node-to-Graph Regression (NGR) model. To specify, the proposed NGR model is composed of two main steps, namely network establishment and network training. In the first step, following the classical researches of image quality estimation that include local distortion measurement followed by pooling for inferring the image quality score, this study captures the local distortion of the input infrared images by stacking up a set of Vision Graph (VSG) blocks to generate one node map, and then conducts the weighted pooling method on the node map to yield the graph output as the estimated quality score. In the second step, for enhancing the model's performance and generalization ability in the network training process, this study implements the node regression with the big data pre-training method to raise the local distortion extraction ability in a broad range of image scenarios and distortion intensities, and then performs the graph regression by using the knowledge distillation method to reduce the over-fitting risk. Using the largest-size infrared image quality evaluation database (I2QED), this study compared the proposed NGR model with three dozen mainstream and state-of-the-art competitors, and results showed that our proposed NGR model achieved the optimal performance. Ke Gu 0001, Hongyan Liu 0004, Yubin Gao, Chen Wang 0019, Lai-Kuan Wong, Weisi Lin, Guangtao Zhai, Wenjun Zhang 0001, Daniel Thalmann |
IEEE Trans. Multim. | 8 |
| 2026 | Age of Semantic Information-Aware Wireless Transmission for Remote Monitoring SystemsabstractSemantic communication is emerging as an effective means of facilitating intelligent and context-aware communication for next-generation communication systems. In this paper, we propose a novel metric called Age of Incorrect Semantics (AoIS) for the transmission of video frames over multiple-input multiple-output (MIMO) channels in a monitoring system. Different from the conventional age-based approaches, we jointly consider the information freshness and the semantic importance, and then formulate a time-averaged AoIS minimization problem by jointly optimizing the semantic actuation indicator, transceiver beamformer, and the semantic symbol design. We first transform the original problem into a low-complexity problem via the Lyapunov optimization. Then, we decompose the transformed problem into multiple subproblems and adopt the alternative optimization (AO) method to solve each subproblem. Specifically, we propose two efficient algorithms, i.e., the successive convex approximation (SCA) algorithm and the low-complexity zero-forcing (ZF) algorithm for optimizing transceiver beamformer. We adopt exhaustive search methods to solve the semantic actuation policy indicator optimization problem and the transmitted semantic symbol design problem. Experimental results demonstrate that our scheme can preserve more than 50% of the original information under the same AoIS compared to the constrained baselines. Xue Han 0003, Biqian Feng, Yongpeng Wu 0001, Xiang-Gen Xia 0001, Wenjun Zhang 0001, Shengli Sun |
IEEE Trans. Wirel. Commun. | 5 |
| 2026 | Semantic Noise-Aided Secure Image Transmission Over MIMO Fading Channels
Xue Han 0003, Biqian Feng, Yongpeng Wu 0001, Yuanwei Liu, Arumugam Nallanathan, Xiang-Gen Xia 0001, Wenjun Zhang 0001 |
IEEE Trans. Wirel. Commun. | 8 |
| 2026 | MambaJSCC: Adaptive Deep Joint Source-Channel Coding With Generalized State Space ModelabstractLightweight and efficient neural network models for deep joint source-channel coding (JSCC) are crucial for semantic communications. In this paper, we propose a novel JSCC architecture, named MambaJSCC, that achieves great performance with low computational and parameter overhead. MambaJSCC utilizes the visual state space model with channel adaptation (VSSM-CA) blocks as its backbone for transmitting images over wireless channels, where the VSSM-CA primarily consists of the generalized state space models (GSSM) and the zero-parameter, zero-computational channel adaptation method (CSI-ReST). We design the GSSM module, leveraging reversible matrix transformations to express generalized scan expanding operations, and theoretically prove that two GSSM modules can effectively capture global information. We discover that GSSM inherently possesses the ability to adapt to channels, a form of endogenous intelligence. Based on this, we design the CSI-ReST method, which injects channel state information (CSI) into the initial state of GSSM to utilize its native response, and into the residual state to mitigate CSI forgetting, enabling effective channel adaptation without introducing additional computational and parameter overhead. Experimental results on different devices, including IoT device JETSON AGX ORIN, show that MambaJSCC not only outperforms existing JSCC methods (e.g., SwinJSCC) across various scenarios but also significantly reduces parameter size, computational overhead, and inference delay. We have released our code and pre-trained models at https://github.com/Wireless3C-SJTU/MambaJSCC, allowing full reproduction of our results. Tong Wu 0003, Zhiyong Chen 0002, Meixia Tao, Xiaodong Xu 0001, Wenjun Zhang 0001, Ping Zhang 0003 |
IEEE Trans. Wirel. Commun. | 6 |
| 2026 | WDMoE: Wireless Distributed Mixture of Experts for Large Language Models
Nan Xue 0007, Zhiyong Chen 0002, Meixia Tao, Xiaodong Xu 0001, Liang Qian, Shuguang Cui, Wenjun Zhang 0001, Ping Zhang 0003 |
IEEE Trans. Wirel. Commun. | 8 |
| 2025 | Mipmap-GS: Let Gaussians Deform with Scale-Specific Mipmap for Anti-Aliasing Renderingabstract3D Gaussian Splatting (3DGS) has attracted great attention in novel view synthesis because of its superior rendering efficiency and high fidelity. However, the trained Gaussians suffer from severe zooming degradation due to non-adjustable representation derived from single-scale training. Though some methods attempt to tackle this problem via post-processing techniques such as selective rendering or filtering techniques towards primitives, the scale-specific information is not involved in Gaussians. In this paper, we propose a unified optimization method to make Gaussians adaptive for arbitrary scales by self-adjusting the primitive properties (e.g., color, shape and size) and distribution (e.g., position). Inspired by the mipmap technique, we design pseudo ground-truth for the target scale and propose a scale-consistency guidance loss to inject scale information into 3D Gaussians. Our method is a plug-in module, applicable for any 3DGS models to solve the zoomin and zoom-out aliasing. Extensive experiments demonstrate the effectiveness of our method. Notably, our method outperforms 3DGS in PSNR by an average of 9.25 dB for zoom-in and 10.40 dB for zoom-out on NeRF Synthetic dataset. Our project website: https://github.com/renaissanceee/Mipmap-GS. Jiezhang Cao, Bingbing Ni, Wenjun Zhang 0001, Kai Zhang 0008, Luc Van Gool |
3DV | 5 |
| 2025 | Neural Block Compression: Variable Bitrates Feature Blocks for Texture RepresentationabstractThe imperative for compression of material textures emerges from the critical demand for high-quality rendering, which necessitates sophisticated textures that, in turn, require substantial storage and memory resources. Thus, low-bitrate compression is crucial, especially in modern games demanding higher texture resolutions. Concurrent methodologies in texture compression predominantly employ a block-based paradigm based on color space, which inevitably leads to representational redundancies and a limited compression scope, particularly at lower bitrates. In the context of mobile devices, bandwidth during texture loading and runtime memory are major bottlenecks, making existing compression algorithms inadequate for high-resolution textures. To mitigate these limitations, we propose a novel multi-resolution texture compression scheme, Neural Block Compression (NBC), developed within the neural feature domain. Our encoding scheme is constructed on a hierarchy of multi-resolution neural feature blocks, and the key ingredient is the variable bitrates quantization scheme. It allocates higher bitrates to higher feature mip-levels and lower bitrates to lower feature mip-levels, thereby extending the concept of block compression from color domain into neural feature domain. Extensive experiments demonstrate the superior texture compression quality achieved by the proposed scheme, especially at low bitrates. Yishun Dou, Xiangzhong Fang, Wenjun Zhang 0001, Bingbing Ni |
AAAI | 5 |
| 2025 | InstantSticker: Realistic Decal Blending via Disentangled Object ReconstructionabstractWe present InstantSticker, a disentangled reconstruction pipeline based on Image-Based Lighting (IBL), which focuses on highly realistic decal blending, simulates stickers attached to the reconstructed surface, and allows for instant editing and real-time rendering. To achieve stereoscopic impression of the decal, we introduce shadow factor into IBL, which can be adaptively optimized during training. This allows the shadow brightness of surfaces to be accurately decomposed rather than baked into the diffuse color, ensuring that the edited texture exhibits authentic shading. To address the issues of warping and blurriness in previous methods, we apply As-Rigid-As-Possible (ARAP) parameterization to pre-unfold a specified area of the mesh and use the local UV mapping combined with a neural texture map to enhance the ability to express high-frequency details in that area. For instant editing, we utilize the Disney BRDF model, explicitly defining material colors with 3-channel diffuse albedo. This enables instant replacement of albedo RGB values during the editing process, avoiding the prolonged optimization required in previous approaches. In our experiment, we introduce the Ratio Variance Warping (RVW) metric to evaluate the local geometric warping of the decal area. Extensive experimental results demonstrate that our method surpasses previous decal blending methods in terms of editing quality, editing speed and rendering speed, achieving the state-of-the-art. Yishun Dou, Ye Chen 0006, Bingbing Ni, Wenjun Zhang 0001 |
AAAI | 8 |
| 2025 | Controllable Distortion-Perception Tradeoff Through Latent Diffusion for Neural Image CompressionabstractNeural image compression often faces a challenging trade-off among rate, distortion and perception. While most existing methods typically focus on either achieving high pixel-level fidelity or optimizing for perceptual metrics, we propose a novel approach that simultaneously addresses both aspects for a fixed neural image codec. Specifically, we introduce a plug-and-play module at the decoder side that leverages a latent diffusion process to transform the decoded features, enhancing either low distortion or high perceptual quality without altering the original image compression codec. Our approach facilitates fusion of original and transformed features without additional training, enabling users to flexibly adjust the balance between distortion and perception during inference. Extensive experimental results demonstrate that our method significantly enhances the pretrained codecs with a wide, adjustable distortion-perception range while maintaining their original compression capabilities. For instance, we can achieve more than 150% improvement in LPIPS-BDRate without sacrificing more than 1 dB in PSNR. Chuqin Zhou, Guo Lu, Jiangchuan Li, Zhengxue Cheng, Li Song 0001, Wenjun Zhang 0001 |
AAAI | 7 |
| 2025 | SinGS: Animatable Single-Image Human Gaussian Splats with Kinematic PriorsabstractDespite significant advances in accurately estimating geometry in contemporary single-image 3D human reconstruction, creating a high-quality, efficient, and animatable 3D avatar remains an open challenge. Two key obstacles persist: incomplete observation and inconsistent 3D priors. To address these challenges, we propose SinGS, aiming to achieve high-quality and efficient animatable 3D avatar reconstruction. At the heart of SinGS are two key components: Kinematic Human Diffusion and Geometry-Preserving 3D Gaussain Splatting. The former is a foundational human model that samples within pose space to generate a highly 3D-consistent and high-quality sequence of human images, inferring unseen viewpoints and providing kinematic priors. The latter is a system that reconstructs a compact, high-quality 3D avatar even under imperfect priors, achieved through a novel semantic Laplacian regularization and a geometry-preserving density control strategy that enable precise and compact assembly of 3D primitives. Extensive experiments demonstrate that SinGS enables lifelike, animatable human reconstructions, maintaining both high quality and inference efficiency (up to 70FPS). Xuanhong Chen, Shunran Jia, Hualiang Wei, Kairui Feng, Yuhan Li 0003, Ang He, Bingbing Ni, Wenjun Zhang 0001 |
CVPR | 12 |
| 2025 | Joint Lossy Compression for a Vector Gaussian Source under Individual Distortion Criteria
Shuao Chen, Junyuan Gao, Yuxuan Shi 0001, Yongpeng Wu 0001, Giuseppe Caire, H. Vincent Poor, Wenjun Zhang 0001 |
GLOBECOM | 7 |
| 2025 | Secure Semantic Image Transmission over Wiretap ChannelsabstractExisting semantic communications have exhibited satisfactory performance in many tasks, but secure image transmission has not been adequately investigated. In this paper, we propose a novel secure semantic image transmission (SSIT) framework over multiple-input single-output (MISO) wiretap channels. To enhance image transmission security for a legitimate semantic user (SU) while interfering with the eavesdropper (Eve), a type of beneficial semantic noise map, which is determined by both source and SU channel states, is produced by a semantic noise-aided conditional variational autoencoder (SN-CVAE) in an unsupervised manner. Furthermore, to improve the secure image reconstruction quality, we propose an efficient transmit beamformer optimization algorithm and leverage the constrained stochastic successive convex approximation (CSSCA) to solve the optimization problem. Numerical results demonstrate that our method effectively protects image information from eavesdroppers while ensuring high-fidelity image reconstruction at the legitimate receiver. Xue Han 0003, Biqian Feng, Yongpeng Wu 0001, Xiang-Gen Xia 0001, Wenjun Zhang 0001 |
GLOBECOM | 5 |
| 2025 | GLDPC Codes Based on Polar Constraints and Their Near-Optimal DecodingabstractIn this work, we introduce the integration of generalized low-density parity-check (GLDPC) codes with short polar component codes, termed GLDPC codes with polar component codes (GLDPC-PC). A recently proposed soft-input soft-output (SISO) decoder for polar-like codes enables effective iterative belief propagation decoding for GLDPC-PC. This SISO decoder after a post-processing exhibits little performance loss to the optimal SISO decoder when all the variable nodes have relatively low degrees. A three-step method is introduced to design protograph-based GLDPC codes. The constructed GLDPC codes are compared with 5G LDPC codes. They exhibit little performance loss in the waterfall region and possess better error floor with less iterations. Binghui Shi, Yongpeng Wu 0001, Yin Xu 0001, Xiqi Gao 0001, Xiaohu You 0001, Wenjun Zhang 0001 |
GLOBECOM | 6 |
| 2025 | Energy Efficiency Maximization for Movable Antenna-Enhanced System Based on Statistical CSI
Xintai Chen, Biqian Feng, Yongpeng Wu 0001, Wenjun Zhang 0001 |
ICC | 4 |
| 2025 | Grant-Free Random Access in Uplink LEO Satellite Communications with OFDMabstractThis paper investigates joint device activity detection and channel estimation for grant-free random access in Lowearth orbit (LEO) satellite communications. We consider uplink communications from multiple single-antenna terrestrial users to a LEO satellite equipped with a uniform planar array of multiple antennas, where orthogonal frequency division multiplexing (OFDM) modulation is adopted. To combat the severe Doppler shift, a transmission scheme is proposed, where the discrete prolate spheroidal basis expansion model (DPS-BEM) is introduced to reduce the number of unknown channel parameters. Then the vector approximate message passing (VAMP) algorithm is employed to approximate the minimum mean square error estimation of the channel, and the Markov random field is combined to capture the channel sparsity. Meanwhile, the expectation-maximization (EM) approach is integrated to learn the hyperparameters in priors. Finally, active devices are detected by calculating energy of the estimated channel. Simulation results demonstrate that the proposed method outperforms conventional algorithms in terms of activity error rate and channel estimation precision. Rui Mao 0020, Yongpeng Wu 0001, Boxiao Shen, Symeon Chatzinotas, Björn Ottersten 0001, Wenjun Zhang 0001 |
ICC | 6 |
| 2025 | Sum Rate Maximization for Movable Antenna-Aided Downlink RSMA SystemsabstractRate splitting multiple access (RSMA) is regarded as a crucial and powerful physical layer (PHY) paradigm for nextgeneration communication systems. Particularly, users employ successive interference cancellation (SIC) to decode part of the interference while treating the remainder as noise. However, conventional RSMA systems rely on fixed-position antenna arrays, limiting their ability to fully exploit spatial diversity. This constraint reduces beamforming gain and significantly impairs RSMA performance. To address this problem, we propose a movable antenna (MA)-aided RSMA scheme that allows the antennas at the base station (BS) to dynamically adjust their positions. Our objective is to maximize the system sum rate of common and private messages by jointly optimizing the MA positions, beamforming matrix, and common rate allocation. To tackle the formulated non-convex problem, we apply fractional programming (FP) and develop an efficient two-stage, coarse-to-fine-grained searching (CFGS) algorithm to obtain high-quality solutions. Numerical results demonstrate that, with optimized antenna adjustments, the MA-enabled system achieves substantial performance and reliability improvements in RSMA over fixedposition antenna setups. Cixiao Zhang, Size Peng, Yin Xu 0001, Qingqing Wu 0001, XiaoWu Ou, Xinghao Guo, Dazhi He, Wenjun Zhang 0001 |
ICC | 8 |
| 2025 | Decentralized Hybrid Precoding for Massive Mu-Mimo IsacabstractIntegrated sensing and communication (ISAC) is a very promising technology designed to provide both high rate communication capabilities and sensing capabilities. However, in Massive Multi User Multiple-Input Multiple-Output (Massive MU MIMO-ISAC) systems, the dense user access creates a serious multi-user interference (MUI) problem, leading to degradation of communication performance. To alleviate this problem, we propose a decentralized baseband processing (DBP) precoding method. We first model the MUI of dense user scenarios with minimizing Cramér-Rao bound (CRB) as an objective function. Hybrid precoding is an attractive ISAC technique, and hybrid precoding using Partially Connected Structures (PCS) can effectively reduce hardware cost and power consumption. We mitigate the MUI between dense users based on Thomlinson-Harashima Precoding (THP). We demonstrate the effectiveness of the proposed method through simulation experiments. Compared with the existing methods, it can effectively improve the communication data rates and energy efficiency in dense user access scenario, and reduce the hardware complexity of Massive MU MIMO-ISAC systems. The experimental results demonstrate the usefulness of our method for improving the MUI problem in ISAC systems for dense user access scenarios. Yin Xu 0001, Dazhi He, Haoyang Li 0004, Yunfeng Guan 0001, Wenjun Zhang 0001 |
ICC | 6 |
| 2025 | Knowledge Distillation for Learned Image Compression
Yunuo Chen 0002, Zezheng Lyu, Guo Lu, Wenjun Zhang 0001 |
ICCV | 7 |
| 2025 | OBI-Bench: Can LMMs Aid in Study of Ancient Script on Oracle Bones?abstractWe introduce OBI-Bench, a holistic benchmark crafted to systematically evaluate large multi-modal models (LMMs) on whole-process oracle bone inscriptions (OBI) processing tasks demanding expert-level domain knowledge and deliberate cognition. OBI-Bench includes 5,523 meticulously collected diverse-sourced images, covering five key domain problems: recognition, rejoining, classification, retrieval, and deciphering. These images span centuries of archaeological findings and years of research by front-line scholars, comprising multi-stage font appearances from excavation to synthesis, such as original oracle bone, inked rubbings, oracle bone fragments, cropped single characters, and handprinted characters. Unlike existing benchmarks, OBI-Bench focuses on advanced visual perception and reasoning with OBI-specific knowledge, challenging LMMs to perform tasks akin to those faced by experts. The evaluation of 6 proprietary LMMs as well as 17 open-source LMMs highlights the substantial challenges and demands posed by OBI-Bench. Even the latest versions of GPT-4o, Gemini 1.5 Pro, and Qwen-VL-Max are still far from public-level humans in some fine-grained perception tasks. However, they perform at a level comparable to untrained humans in deciphering tasks, indicating remarkable capabilities in offering new interpretative perspectives and generating creative guesses. We hope OBI-Bench can facilitate the community to develop domain-specific multi-modal foundation models towards ancient language research and delve deeper to discover and enhance these untapped potentials of LMMs. Zijian Chen 0001, Tingzhu Chen, Wenjun Zhang 0001, Guangtao Zhai |
ICLR | 3 |
| 2025 | MultiEgo: A Multi-View Egocentric Video Dataset for 4D Scene ReconstructionabstractMulti-view egocentric dynamic scene reconstruction holds significant research value for applications in holographic documentation of social interactions. However, existing reconstruction datasets focus on static multi-view or single-egocentric view setups, lacking multi-view egocentric datasets for dynamic scene reconstruction. Therefore, we present MultiEgo, the first multi-view egocentric dataset for 4D dynamic scene reconstruction. The dataset comprises five canonical social interaction scenes: meetings, performances, and a presentation. Each scene provides five authentic egocentric videos captured by participants wearing AR glasses. We design a hardware-based data acquisition system and processing pipeline, achieving sub-millisecond temporal synchronization across views, coupled with accurate pose annotations. Experiment validation demonstrates the practical utility and effectiveness of our dataset for free-viewpoint video (FVV) applications, establishing MultiEgo as a foundational resource for advancing multi-view egocentric dynamic scene reconstruction research. Bate Li, Houqiang Zhong, Zhengxue Cheng, Qiang Hu 0003, Qiang Wang 0061, Li Song 0001, Wenjun Zhang 0001 |
ACM Multimedia | 7 |
| 2025 | H3D-DGS: Exploring Heterogeneous 3D Motion Representation for Deformable 3D Gaussian SplattingabstractDynamic scene reconstruction poses a persistent challenge in 3D vision. Deformable 3D Gaussian Splatting has emerged as an effective method for this task, offering real-time rendering and high visual fidelity.
This approach decomposes a dynamic scene into a static representation in a canonical space and time-varying scene motion.
Scene motion is defined as the collective movement of all Gaussian points, and for compactness, existing approaches commonly adopt implicit neural fields or sparse control points.
However, these methods predominantly rely on gradient-based optimization for all motion information. Due to the high degree of freedom, they struggle to converge on real-world datasets exhibiting complex motion.
To preserve the compactness of motion representation and address convergence challenges, this paper proposes heterogeneous 3D control points, termed \textbf{H3D control points}, whose attributes are obtained using a hybrid strategy combining optical flow back-projection and gradient-based methods.
This design decouples directly observable motion components from those that are geometrically occluded.
Specifically, components of 3D motion that project onto the image plane are directly acquired via optical flow back projection, while unobservable portions are refined through gradient-based optimization.
Experiments on the Neu3DV and CMU-Panoptic datasets demonstrate that our method achieves superior performance over state-of-the-art deformable 3D Gaussian splatting techniques. Remarkably, our method converges within just 100 iterations and achieves a per-frame processing speed of 2 seconds on a single NVIDIA RTX 4070 GPU. Yunuo Chen 0002, Guo Lu, Cheems Wang, Qunshan Gu, Rong Xie 0004, Li Song 0001, Wenjun Zhang 0001 |
NeurIPS | 8 |
| 2025 | 4DGCPro: Efficient Hierarchical 4D Gaussian Compression for Progressive Volumetric Video StreamingabstractAchieving seamless viewing of high-fidelity volumetric video, comparable to 2D video experiences, remains an open challenge. Existing volumetric video compression methods either lack the flexibility to adjust quality and bitrate within a single model for efficient streaming across diverse networks and devices, or struggle with real-time decoding and rendering on lightweight mobile platforms. To address these challenges, we introduce 4DGCPro, a novel hierarchical 4D Gaussian compression framework that facilitates real-time mobile decoding and high-quality rendering via progressive volumetric video streaming in a single bitstream. Specifically, we propose a perceptually-weighted and compression-friendly hierarchical 4D Gaussian representation with motion-aware adaptive grouping to reduce temporal redundancy, preserve coherence, and enable scalable multi-level detail streaming. Furthermore, we present an end-to-end entropy-optimized training scheme, which incorporates layer-wise rate-distortion (RD) supervision and attribute-specific entropy modeling for efficient bitstream generation. Extensive experiments show that 4DGCPro enables flexible quality and variable bitrate within a single model, achieving real-time decoding and rendering on mobile devices while outperforming existing methods in RD performance across multiple datasets. Zihan Zheng, Zhenlong Wu, Houqiang Zhong, Yuan Tian 0017, Lan Xu 0003, Jiangchao Yao, Xiaoyun Zhang 0001, Qiang Hu 0003, Wenjun Zhang 0001 |
NeurIPS | 10 |
| 2025 | Two Birds with One Stone: Multi-Task Semantic Communications Systems Over Relay ChannelabstractIn this paper, we propose a novel multi-task, multi-link relay semantic communications (MTML-RSC) scheme that enables the destination node to simultaneously perform image reconstruction and classification with one transmission from the source node. In the MTML-RSC scheme, the source node broadcasts a signal using semantic communications, and the relay node forwards the signal to the destination. We analyze the coupling relationship between the two tasks and the two links (source-to-relay and source-to-destination) and design a semantic-focused forward method for the relay node, where it selectively forwards only the semantics of the relevant class while ignoring others. At the destination, the node combines signals from both the source node and the relay node to perform classification, and then uses the classification result to assist in decoding the signal from the relay node for image reconstructing. Experimental results demonstrate that the proposed MTML-RSC scheme achieves significant performance gains, e.g., 1.73 dB improvement in peak-signal-to-noise ratio (PSNR) for image reconstruction and increasing the accuracy from 64.89% to 70.31 % for classification. Tong Wu 0003, Zhiyong Chen 0002, Yin Xu 0001, Meixia Tao, Wenjun Zhang 0001 |
WCNC | 6 |
| 2025 | Fluid Antenna Grouping Index Modulation Design for MIMO SystemsabstractThe fluid antenna (FA)-enabled multiple-input multiple-output (MIMO) system based on index modulation (IM), referred to as FA-IM, significantly enhances spectral efficiency (SE) compared to the conventional FA-assisted MIMO system. To improve the performance in addressing the high spatial correlations between multiple activated ports, this paper proposes an innovative FA grouping-based IM (FAG-IM) system. Specifically, considering the characteristics of the FA two-dimensional (2D) surface structure and the spatially correlated channel model in FA-assisted MIMO systems, a block grouping method is adopted, where adjacent ports are assigned to the same group. Consequently, different groups independently perform port index selection and constellation symbol mapping, with only one port being activated within each group during each transmission interval. Then, a closed-form average bit error probability (ABEP) upper bound is derived for the proposed system. Numerical results show that, compared to state-of-the-art systems, the FAG-IM system consistently achieves substantial performance gains. Xinghao Guo, Yin Xu 0001, Dazhi He, Cixiao Zhang, Wenjun Zhang 0001, Yiyan Wu 0001 |
WCNC | 5 |
| 2025 | Soft-Output Fast Successive-Cancellation List Decoder for Polar CodesabstractThe soft-output successive cancellation list (SO-SCL) decoder provides a methodology for estimating the a-posteriori probability log-likelihood ratios by only leveraging the conventional SCL decoder for polar codes. However, the sequential nature of SCL decoding leads to a high decoding latency for the SO-SCL decoder. In this paper, we propose a soft-output fast SCL (SO-FSCL) decoder by incorporating node-based fast decoding into the SO-SCL framework. Simulation results demonstrate that the proposed SO-FSCL decoder significantly reduces the decoding latency without loss of performance compared with the SO-SCL decoder. Yongpeng Wu 0001, Yin Xu 0001, Xiaohu You 0001, Xiqi Gao 0001, Wenjun Zhang 0001 |
WCNC | 6 |
| 2025 | Terahertz aerospace communications: enabling technologies and future directions
Weijun Gao 0001, Chong Han 0001, Yuanzhi He, Wenjun Zhang 0001 |
Sci. China Inf. Sci. | 5 |
| 2025 | Distributed Deep Reinforcement Learning-Based Gradient Quantization for Federated Learning Enabled Vehicle Edge ComputingabstractFederated learning (FL) can protect the privacy of the vehicles in vehicle edge computing (VEC) to a certain extent through sharing the gradients of vehicles’ local models instead of the local data. The gradients of vehicles’ local models are usually large for the vehicular artificial intelligence (AI) applications, thus transmitting such large gradients would cause large per-round latency. Gradient quantization has been proposed as one effective approach to reduce the per-round latency in FL enabled VEC through compressing gradients and reducing the number of bits, i.e., the quantization level, to transmit gradients. The selection of quantization level and thresholds determines the quantization error (QE), which further affects the model accuracy and training time. To do so, the total training time and QE become two key metrics for the FL enabled VEC. It is critical to jointly optimize the total training time and QE for the FL enabled VEC. However, the time-varying channel condition causes more challenges to solve this problem. In this article, we propose a distributed deep reinforcement learning (DRL)-based quantization level allocation scheme to optimize the long-term reward in terms of the total training time and QE. Extensive simulations identify the optimal weighted factors between the total training time and QE, and demonstrate the feasibility and effectiveness of the proposed scheme. Wenjun Zhang 0001, Qiong Wu 0002, Pingyi Fan, Qiang Fan 0002, Jiangzhou Wang, Khaled Ben Letaief |
IEEE Internet Things J. | 2 |
| 2025 | Massive MIMO-OTFS-Based Random Access for Cooperative LEO Satellite ConstellationsabstractThis paper investigates joint device identification, channel estimation, and symbol detection for cooperative multi-satellite-enhanced random access, where orthogonal time-frequency space modulation with the large antenna array is utilized to combat the dynamics of the terrestrial-satellite links (TSLs). We introduce the generalized complex exponential basis expansion model to parameterize TSLs, thereby reducing the pilot overhead. By exploiting the block sparsity of the TSLs in the angular domain, a message passing algorithm is designed for initial channel estimation. Subsequently, we examine two cooperative modes to leverage the spatial diversity within satellite constellations: the centralized mode, where computations are performed at a high-power central server, and the distributed mode, where computations are offloaded to edge satellites with minimal signaling overhead. Specifically, in the centralized mode, device identification is achieved by aggregating backhaul information from edge satellites, and channel estimation and symbol detection are jointly enhanced through a structured approximate expectation propagation (AEP) algorithm. In the distributed mode, edge satellites share channel information and exchange soft information about data symbols, leading to a distributed version of AEP. The introduced basis expansion model for TSLs enables the efficient implementation of both centralized and distributed algorithms via fast Fourier transform. Simulation results demonstrate that proposed schemes significantly outperform conventional algorithms in terms of the activity error rate, the normalized mean squared error, and the symbol error rate. Notably, the distributed mode achieves performance comparable to the centralized mode with only two exchanges of soft information about data symbols within the constellation. Boxiao Shen, Yongpeng Wu 0001, Shiqi Gong, Heng Liu 0007, Björn Ottersten 0001, Wenjun Zhang 0001 |
IEEE J. Sel. Areas Commun. | 6 |
| 2025 | Probabilistic Shaped Multilevel Polar Coding for Wiretap ChannelabstractA wiretap channel is served as the fundamental model of physical layer security techniques, where the secrecy capacity of the Gaussian wiretap channel is proven to be achieved by Gaussian input. However, there remains a gap between the Gaussian secrecy capacity and the secrecy rate with conventional uniformly distributed discrete constellation input, e.g. amplitude shift keying (ASK) and quadrature amplitude modulation (QAM). In this paper, we propose a probabilistic shaped multilevel polar coding scheme to bridge the gap. Specifically, the input distribution optimization problem for maximizing the secrecy rate with ASK/QAM input is solved. Numerical results show that the resulting sub-optimal solution can still approach the Gaussian secrecy capacity. Then, we investigate the polarization of multilevel polar codes for the asymmetric discrete memoryless wiretap channel, and thus propose a multilevel polar coding scheme integration with probabilistic shaping. It is proved that the scheme can achieve the secrecy capacity of the Gaussian wiretap channel with discrete constellation input, and satisfies the reliability condition and weak security condition. A security-oriented polar code construction method to natively satisfies the leakage-based security condition is also investigated. Simulation results show that the proposed scheme achieves more efficient and secure transmission than the uniform constellation input case over both the Gaussian wiretap channel and the Rayleigh fading wiretap channel. Yongpeng Wu 0001, Peihong Yuan, Chengshan Xiao, Xiang-Gen Xia 0001, Wenjun Zhang 0001 |
IEEE J. Sel. Areas Commun. | 6 |
| 2025 | RWZC: A Model-Driven Approach for Learning-Based Robust Wyner-Ziv CodingabstractIn this paper, a novel learning-based Wyner-Ziv coding framework is considered under a distributed image transmission scenario, where the correlated source is only available at the receiver. Unlike other learnable frameworks, our approach demonstrates robustness to non-stationary source correlation, where the overlapping information between image pairs varies. Specifically, we first model the affine relationship between correlated images and leverage this model for learnable mask generation and rate-adaptive joint source-channel coding. Moreover, we also provide a warping-prediction network to remove the distortion from channel interference and affine transform. Intuitively, the observed performance improvement is largely due to focusing on the simple geometric relationship, rather than the complex joint distribution between the sources. Numerical results show that our framework achieves a 1.5 dB gain in PSNR and a 0.2 improvement in MS-SSIM, along with a significant superiority in perceptual metric, compared to state-of-the-art methods when applied to real-world samples with non-stationary correlations. Yuxuan Shi 0001, Shuo Shao 0001, Yongpeng Wu 0001, Wenjun Zhang 0001, Mérouane Debbah |
IEEE J. Sel. Areas Commun. | 4 |
| 2025 | Fluid Antenna Index Modulation for MIMO Systems: Robust Transmission and Low-Complexity Detection
Xinghao Guo, Yin Xu 0001, Dazhi He, Cixiao Zhang, Hanjiang Hong, Kai-Kit Wong, Wenjun Zhang 0001, Yiyan Wu 0001 |
IEEE Trans. Commun. | 7 |
| 2025 | SCSC: A Novel Standards-Compatible Semantic Communication Framework for Image TransmissionabstractJoint source-channel coding (JSCC) is a promising paradigm for next-generation communication systems, particularly in challenging transmission environments. In this paper, we propose a novel standard-compatible JSCC framework for the transmission of images over multiple-input multiple-output (MIMO) channels. Different from the existing end-to-end AI-based DeepJSCC schemes, our framework consists of learnable modules that enable communication using conventional separate source and channel codes (SSCC), which makes it amenable for easy deployment on legacy systems. Specifically, the learnable modules involve a preprocessing-empowered network (PPEN) for preserving essential semantic information, and a precoder & combiner-enhanced network (PCEN) for efficient transmission over a resource-constrained MIMO channel. We treat existing compression and channel coding modules as non-trainable blocks. Since the parameters of these modules are non-differentiable, we employ a proxy network that mimics their operations when training the learnable modules. Numerical results demonstrate that our scheme can save more than 29% of the channel bandwidth, and requires lower complexity compared to the constrained baselines. We also show its generalization capability to unseen datasets and tasks through extensive experiments. Xue Han 0003, Yongpeng Wu 0001, Zhen Gao 0001, Biqian Feng, Yuxuan Shi 0001, Deniz Gündüz, Wenjun Zhang 0001 |
IEEE Trans. Commun. | 7 |
| 2025 | Time-Smooth Wireless Transmission of Probabilistic Slicing VR 360 Video in MISO-OFDM SystemsabstractThe multiple-input and single-output (MISO)-orthogonal frequency-division multiplexing (OFDM) systems afford low latency and high reliability for virtual reality (VR) 360 video in multi-user scenarios. Motivated by the goal of maintaining time-smoothness while holding acceptably low complexity, a crucial factor in VR video transmission, we conduct a comprehensive study that integrates the characteristics of VR video with the strategies for subcarrier assignment and power allocation. By analyzing the pre-transmitted tile-segments, the missing tile-segments, and the video frame structure, we propose two probabilistic slicing schemes (PSPs) to minimize the size of required tile-segments of VR video scenes. In time-smoothness maximization, the desired discrete encoding rate set, discrete subcarrier assignment, continuous power allocation, and fixed total power constraint make it a challenging mixed-integer nonlinear programming (MINLP) problem. Unlike the straightforward relaxation-recovery method, we firstly prove that a near-optimal recovered encoding rate is the discrete value closest to the optimal relaxed-continuous encoding rate. We then propose a Two-step Encoding Rate Maximization (TERM) method, including the relaxed-continuous sum-rate maximization and the discrete encoding rate recovery, to achieve the near-optimal subcarrier assignment and the power allocation with low complexity. Simulation results on real-world VR video dataset validate that the two PSPs can effectively minimize the number of transmitted tile-segments. The proposed TERM with PSPs can maintain time-smoothness of VR 360 video with an acceptably low level of complexity in MISO-OFDM systems. Guangtao Zhai, Yongpeng Wu 0001, Xiongkuo Min, Biqian Feng, Yucheng Zhu, Wenjun Zhang 0001 |
IEEE Trans. Commun. | 7 |
| 2025 | Joint Luminance-Chrominance Learning for Image DebandingabstractBanding is a visually annoying artifact that frequently occurs along the chain of video acquisition, production, distribution, and display, showing a significant need for improvement in many fields. Thus far, efforts on banding removal are mainly knowledge-driven or merely learning on RGB space, which is either limited by domain knowledge or lacks the consideration for banding in chrominance channels. In this work, we propose a unified deep neural network that explicitly disentangles the luminance and chrominance channels, and simultaneously recovers intensity gradients and color discontinuity from detection-free measurement in an end-to-end manner. Our debanding model is comprised of a luminance restoration network (LR-Net) and a chrominance restoration network (CR-Net). Each of them follows an encoder-decoder architecture, where a cascade of residual blocks is employed to exploit hierarchical non-local features in spatial dimensions for more powerful feature representation. Moreover, we investigate the characteristics of banding artifacts and apply specific loss functions to guide the debanding in different channels, thus boosting the restoration performance. Both qualitative and quantitative experiments show that our model significantly surpasses the existing method in terms of all 7 metrics. Ultimately, our network trained on simulated data exhibits good adaptiveness under various compression scenarios, which further demonstrates the effectiveness of the proposed model. Zijian Chen 0001, Wei Sun 0029, Jun Jia, Ru Huang 0002, Fangfang Lu, Ying Chen 0011, Xiongkuo Min, Guangtao Zhai, Wenjun Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 9 |
| 2025 | Study of Subjective and Objective Naturalness Assessment of AI-Generated ImagesabstractThe proliferation of Artificial Intelligence-Generated Images (AIGIs) has greatly expanded the Image Naturalness Assessment (INA) problem. Different from early definitions that mainly focus on tone-mapped images with limited distortions (e.g., exposure, contrast, and color reproduction), INA on AI-generated images is especially challenging as it owns more diverse contents and could be affected by factors from multiple perspectives, including low-level technical distortions and high-level rationality distortions. In this paper, we take the first step to benchmark and assess the visual naturalness of AI-generated images. First, we construct the AI-Generated Image Naturalness (AGIN) dataset by conducting a large-scale subjective study to collect human opinions on the overall naturalness as well as perceptions from the technical quality and rationality perspectives. AGIN verifies several insights for the first time that naturalness is universally and disparately affected by both technical and rational distortions, while its manifestations vary with different generation tasks. Second, to automatically assess the naturalness of AIGIs that align with human opinions, we propose the Joint Objective Image Naturalness evaluaTor (JOINT). Specifically, JOINT imitates human reasoning in naturalness evaluation by jointly learning technical and rationality features with several specific designs to guide model behavior from respective perspectives. Experiments demonstrate that JOINT significantly outperforms existing methods for providing more subjectively consistent results on naturalness assessment. The dataset can be accessed athttps://github.com/zijianchen98/AGIN. Zijian Chen 0001, Wei Sun 0029, Haoning Wu 0001, Jun Jia, Ru Huang 0002, Xiongkuo Min, Guangtao Zhai, Wenjun Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 9 |
| 2025 | Perceptual Information Fidelity for Quality Estimation of Industrial ImagesabstractDepending on high quality images, industrial vision technologies can basically oversee all the industrial production processes, such as workpiece processing and assembly automation, which play a highly significant role in promoting detection automation and production capacity in assembly lines. Unlike the natural scene images which consist of richer colors and natural lines, industrial images that cover complex industrial goods and equipment are made up of fewer colors, more regular shapes, massive graphic elements, etc., causing existing image processing methods for quality estimation, enhancement and monitoring to fail. Human beings usually play the part of the final receiver of an industrial image, so in the researches of image quality estimation, it is necessary to take the perception process of human eyes and brain to the input images into consideration. On this basis, we in this paper propose a novel perceptual information fidelity based image quality estimation model, abbreviated as PIF. Particularly, we first introduce a visual-cell low-pass filter and an optical-nerve noise model, which are separately inspired by the two processes: one is that an image in the form of optical signals arrives at the retina through the eye’s optical system to form the stimuli; the other is that the aforesaid stimuli in the form of electrical signals transfer to the human brain through the optical nerve. Second, we construct a novel image content-aware adjustor to optimize the above visual-cell low-pass filter and optical-nerve noise model. Third, we compare the two quantities of the information that is present in the clean image and how much of the information can be extracted from the lossy image to generate the overall quality score. Experiments on the two large-size industrial image quality databases demonstrate the excellent performance achieved by our proposed PIF model, with a remarkable performance gain over the existing state-of-the-art competitors. Ke Gu 0001, Hongyan Liu 0004, Junfei Qiao 0001, Guangtao Zhai, Wenjun Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | Exploring Rich Subjective Quality Information for Image Quality Assessment in the WildabstractTraditional in the wild image quality assessment (IQA) models are generally trained with the quality labels of mean opinion score (MOS), while missing the rich subjective quality information contained in the quality ratings, for example, the standard deviation of opinion scores (SOS) or even distribution of opinion scores (DOS). In this paper, we propose a novel IQA method namedRichIQAto explore the rich subjective rating information beyond MOS to predict image quality in the wild. RichIQA is characterized by two key novel designs: 1) a three-stage image quality prediction network which exploits the powerful feature representation capability of the Convolutional vision Transformer (CvT) and mimics the short-term and long-term memory mechanisms of human brain; 2) a multi-label training strategy in which rich subjective quality information like MOS, SOS and DOS are concurrently used to train the quality prediction network. Powered by these two novel designs, RichIQA is able to predict the image quality in terms of a distribution, from which the mean image quality can be subsequently obtained. Extensive experimental results verify that the three-stage network is tailored to predict rich quality information, while the multi-label training strategy can fully exploit the potentials within subjective quality rating and enhance the prediction performance and generalizability of the network. RichIQA outperforms state-of-the-art competitors on multiple large-scale in the wild IQA databases with rich subjective rating labels. The code of RichIQA will be made publicly available on GitHub. Xiongkuo Min, Yuqin Cao, Guangtao Zhai, Wenjun Zhang 0001, Huifang Sun, Chang Wen Chen |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Mesh2Animation: Unsupervised Animating for Quadruped 3D ObjectsabstractAnimating quadruped 3D objects, such as chairs and tables, typically involves three steps in the traditional computer graphics pipeline: Rigging, Skinning, and Retargeting. Commonly, prevailing methods for each specific step are conceived in isolation. For rigging and skinning steps, optimization-based methods are typically used, but these approaches tend to be slow and susceptible to variations in 3D mesh surfaces. For the retargeting step, the obtained results often fall short of expectations, especially when dealing with dissimilar source and target skeletons, leading to issues like joint twisting. The devised procedure is also time-intensive, resulting in a complex final pipeline. To this end, we present a unified framework, termed Mesh2Animation, providing an end-to-end solution to these challenges. In Mesh2Animation, a learning-based method is proposed for quadruped 3D skeleton estimation. We introduce both skeleton-level and mesh-level loss, allowing the rigging, skinning, and retargeting steps to be optimized simultaneously. Specifically, a general predicted estimation from the rigging step initializes the skeleton, making the skinning step faster and more accurate, which in turn leads to better results in the retargeting step. Finally, the rigging, skinning and retargeting processes are optimized simultaneously under static and temporal constraints. Additionally, we can construct a novel animating dataset termed ShapeNet2Animation (SN2Animation) based on the proposed method, which shows potential application for pose transfer. Qualitative and quantitative results on SN2Animation, ShapeNet, Object3D and ModelNet10 datasets for animation demonstrate that our method achieves competitive performance and shows promising generalization ability on quadruped 3D objects. Our project is available athttps://sites.google.com/view/mesh2animation. Zhenbo Yu, Jinxian Liu, Zefan Li, Bingbing Ni, Wenjun Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2025 | SSP-IR: Semantic and Structure Priors for Diffusion-Based Realistic Image RestorationabstractRealistic image restoration is a crucial task in computer vision, and diffusion-based models for image restoration have garnered significant attention due to their ability to produce realistic results. Restoration can be seen as a controllable generation conditioning on priors. However, due to the severity of image degradation, existing diffusion-based restoration methods cannot fully exploit priors from low-quality images and still have many challenges in perceptual quality, semantic fidelity, and structure accuracy. Based on the challenges, we introduce a novel image restoration method, SSP-IR. Our approach aims to fully exploit semantic and structure priors from low-quality images to guide the diffusion model in generating semantically faithful and structurally accurate natural restoration results. Specifically, we integrate the visual comprehension capabilities of Multimodal Large Language Models (explicit) and the visual representations of the original image (implicit) to acquire accurate semantic prior. To extract degradation-independent structure prior, we introduce a Processor with RGB and FFT constraints to extract structure prior from the low-quality images, guiding the diffusion model and preventing the generation of unreasonable artifacts. Lastly, we employ a multi-level attention mechanism to integrate the acquired semantic and structure priors. The qualitative and quantitative results demonstrate that our method outperforms other state-of-the-art methods overall on both synthetic and real-world datasets. Our project page ishttps://zyhrainbow.github.io/projects/SSP-IR. Hengsheng Zhang, Zhengxue Cheng, Rong Xie 0004, Li Song 0001, Wenjun Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | MISC: Ultra-Low Bitrate Image Semantic Compression Driven by Large Multimodal ModelabstractWith the evolution of storage and communication protocols, ultra-low bitrate image compression has become a highly demanding topic. However, all existing compression algorithms must sacrifice either consistency with the ground truth or perceptual quality at ultra-low bitrate. During recent years, the rapid development of the Large Multimodal Model (LMM) has made it possible to balance these two goals. To solve this problem, this paper proposes a method called Multimodal Image Semantic Compression (MISC), which consists of an LMM encoder for extracting the semantic information of the image, a map encoder to locate the region corresponding to the semantic, an image encoder generates an extremely compressed bitstream, and a decoder reconstructs the image based on the above information. Experimental results show that our proposed MISC is suitable for compressing both traditional Natural Sense Images (NSIs) and emerging AI-Generated Images (AIGIs) content. It can achieve optimal consistency and perception results while saving 50% bitrate, which has strong potential applications in the next generation of storage and communication. The code will be released on https://github.com/lcysyzxdxc/MISC. Chunyi Li 0001, Guo Lu, Donghui Feng 0003, Haoning Wu 0001, Xiaohong Liu 0001, Guangtao Zhai, Weisi Lin, Wenjun Zhang 0001 |
IEEE Trans. Image Process. | 9 |
| 2025 | Unsourced Random Access in MIMO Quasi-Static Rayleigh Fading Channels: Finite Blocklength and Scaling Law Analyses
Junyuan Gao, Yongpeng Wu 0001, Giuseppe Caire, Wei Yang 0001, H. Vincent Poor, Wenjun Zhang 0001 |
IEEE Trans. Inf. Theory | 6 |
| 2025 | Air Pollution Monitoring by Integrating Local and Global Information in Self-Adaptive Multiscale Transform DomainabstractThis paper proposed a novel image-based air pollution monitor (IAPM) by incorporating local and global information in the self-adaptive multiscale transform domain, so as to achieve the timely and effective leakage detection of typical air pollutants from a single image. To be specific, this paper first developed a screen-shaped module according to two significant findings in visual neuroscience, which include the high sensitivity of human eyes to horizontal and vertical stimuli and the center-surround inhibition, by designing and fusing the square module, horizontal strip module and vertical strip module parallelly for simulating the behaviour of human eyes to extract local features. Second, the learnable weights and proportional mapping were applied to incorporate the screen-shaped module and lightweight vision transformer as backbone, towards more richly exploiting and fusing local and global information just as the way a brain perceives external stimuli. Third, a new self-adaptive multiscale transform domain method was devised based on two motivations from the visual characteristics of multiscale perception and the brain characteristics of self-adaptive domain transform to modify the backbone by using the operations of pooling and pointwise convolution. Extensive experiments implemented on the datasets of carbon particulate matters and ethylene leakage confirmed the superior monitoring performance of the proposed IAPM model beyond the state-of-the-art (SOTA) peers by an accuracy gain of about 4%. Furthermore, the proposed IAPM model only required 0.089 GFLOPs and 0.15 million model parameters, remarkably outperforming SOTA competitors in computational efficiency and storage resources. Ke Gu 0001, Hongyan Liu 0004, Bo Liu 0024, Junfei Qiao 0001, Weisi Lin, Wenjun Zhang 0001 |
IEEE Trans. Multim. | 7 |
| 2025 | Downlink OFDM-FAMA in 5G-NR SystemsabstractFluid antenna multiple access (FAMA), enabled by the fluid antenna system (FAS), offers a new and straightforward solution to massive connectivity. Previous results on FAMA were primarily based on narrowband channels. This paper studies the adoption of FAMA within the fifth-generation (5G) orthogonal frequency division multiplexing (OFDM) framework, referred to as OFDM-FAMA, and evaluate its performance in broadband multipath channels. We first design the OFDM-FAMA system, taking into account 5G channel coding and OFDM modulation. Then the system’s achievable rate is analyzed, and an algorithm to approximate the FAS configuration at each user is proposed based on the rate. Extensive link-level simulation results reveal that OFDM-FAMA can significantly improve the multiplexing gain over the OFDM system with fixed-position antenna (FPA) users, especially when robust channel coding is applied and the number of radio-frequency (RF) chains at each user is small. Hanjiang Hong, Kai-Kit Wong, Hao Xu 0003, Yin Xu 0001, Hyundong Shin, Ross Murch, Dazhi He, Wenjun Zhang 0001 |
IEEE Trans. Wirel. Commun. | 8 |
| 2025 | Addressing the Curse of Scenario and Task Generalization in AI-6G: A Multi-Modal ParadigmabstractExisting works on machine learning (ML)-empowered wireless communication primarily focus on monolithic scenarios and single tasks. However, with the blooming growth of communication task classes coupled with various task requirements in future 6G systems, this working pattern is obviously unsustainable. Therefore, identifying a groundbreaking paradigm that enables a universal model to solve multiple tasks in the physical layer within diverse scenarios is crucial for future system evolution. This paper aims to fundamentally address the curse of ML model generalization across diverse scenarios and tasks by unleashing multi-modal feature integration capabilities in future systems. Given the universality of electromagnetic propagation theory, the communication process is determined by the scattering environment, which can be more comprehensively characterized by cross-modal perception, thus providing sufficient information for all communication tasks across varied environments. This fact motivates us to propose a transformative two-stage multi-modal pre-training and downstream task adaptation paradigm. In the pre-training stage, we introduce a multi-modal two-tower model and a corresponding contrastive learning method to integrate the explicit description of the scattering environment and implicit channel state information (CSI) into a universal representation, which encapsulates rich high-level knowledge and can be leveraged for all downstream tasks in different scenarios. Additionally, we present two specially designed model structures to enhance the interaction of communication modalities. In the second stage, based on the frozen pre-trained model, we propose a direct method and a pluggable method for flexible and low-cost task adaptation. Experimental results demonstrate that our proposed approach significantly outperforms benchmarks in both task performance and tuning parameter size for exemplary sub-tasks in unseen scenarios. Tianyu Jiao, Zhuoran Xiao, Yin Xu 0001, Chenhui Ye, Zhiyong Chen 0002, Liyu Cai, Dazhi He, Yunfeng Guan 0001, Guangyi Liu 0001, Wenjun Zhang 0001 |
IEEE Trans. Wirel. Commun. | 12 |
| 2025 | Asynchronous MIMO-OFDM Massive Unsourced Random Access With Codeword CollisionsabstractThis paper investigates asynchronous multiple-input multiple-output (MIMO) massive unsourced random access (URA) in an orthogonal frequency division multiplexing (OFDM) system over frequency-selective fading channels, with the presence of both timing and carrier frequency offsets (TO and CFO) and non-negligible codeword collisions. The proposed coding framework segregates the data into two components, namely, preamble and coding parts, with the former being tree-coded and the latter LDPC-coded. By leveraging the dual sparsity of the equivalent channel across both codeword and delay domains (CD and DD), we develop a message-passing-based sparse Bayesian learning algorithm, combined with belief propagation and mean field, to iteratively estimate DD channel responses, TO, and delay profiles. Furthermore, by jointly leveraging the observations among multiple slots, we establish a novel graph-based algorithm to iteratively separate the superimposed channels and compensate for the phase rotations. Additionally, the proposed algorithm is applied to the flat fading scenario to estimate both TO and CFO, where the channel and offset estimation is enhanced by leveraging the geometric characteristics of the signal constellation. Extensive simulations reveal that the proposed algorithm achieves superior performance and substantial complexity reduction in both channel and offset estimation compared to the codebook enlarging-based counterparts, and enhanced data recovery performances compared to state-of-the-art URA schemes. Tianya Li, Yongpeng Wu 0001, Junyuan Gao, Wenjun Zhang 0001, Xiang-Gen Xia 0001, Derrick Wing Kwan Ng, Chengshan Xiao |
IEEE Trans. Wirel. Commun. | 4 |
| 2025 | Low-PAPR Pilot Arrangement and Iterative Channel Estimation for OTFSabstractOrthogonal time frequency space (OTFS) modulation outperforms orthogonal frequency division multiplexing (OFDM) in high-mobility scenarios with doubly selective channels. However, the existing channel estimation schemes for OTFS usually rely on high-power pilots, which cause the issue of high peak-to-average power ratio (PAPR). In this paper, a channel estimation scheme for OTFS is designed, utilizing the Zadoff-Chu (ZC) sequence as the pilot without any guard symbols to reduce the PAPR effectively. Furthermore, an iterative ZC-sequence-based estimation algorithm is proposed. It can accurately estimate channels with integer and fractional Doppler shifts, irrespective of whether ideal or rectangular waveforms are employed. The proposed scheme performs the channel estimation using a correlation-based method. The correlation’s interference, caused by the data, pilot, and noise, is mitigated by performing channel estimation and data detection alternately. Simulation results show that it has a significantly lower PAPR in the time domain and superior channel estimation performance with both integer and fractional Doppler. Tianyao Ma, Yin Xu 0001, XiaoWu Ou, Haoyang Li 0004, Dazhi He, Wenjun Zhang 0001 |
IEEE Trans. Wirel. Commun. | 6 |
| 2025 | LEO Satellite-Enabled Random Access With Large Differential Delay and Doppler ShiftabstractThis paper investigates joint device identification, channel estimation, and symbol detection for LEO satellite-enabled grant-free random access systems, specifically targeting scenarios where remote Internet-of-Things (IoT) devices operate without global navigation satellite system (GNSS) assistance. Considering the constrained power consumption of these devices, the large differential delay and Doppler shift are handled at the satellite receiver. We firstly propose a spreading-based multi-frame transmission scheme with orthogonal time-frequency space (OTFS) modulation to mitigate the doubly dispersive effect in time and frequency, and then analyze the input-output relationship of the system. Next, we propose a receiver structure based on three modules: a linear module for identifying active devices that leverages the generalized approximate message passing algorithm to eliminate inter-user and inter-carrier interference; a non-linear module that employs the message passing algorithm to jointly estimate the channel and detect the transmitted symbols; and a third module that aims to exploit the three dimensional block channel sparsity in the delay-Doppler-angle domain. Soft information is exchanged among the three modules by careful message scheduling. Furthermore, the expectation-maximization algorithm is integrated to adjust phase rotation caused by the fractional Doppler and to learn the hyperparameters in the priors. Finally, the convolutional neural network is incorporated to enhance the symbol detection. Simulation results demonstrate that the proposed transmission scheme boosts the system performance, and the designed algorithms outperform the conventional methods significantly in terms of the device identification, channel estimation, and symbol detection. Boxiao Shen, Yongpeng Wu 0001, Wenjun Zhang 0001, Symeon Chatzinotas, Björn Ottersten 0001 |
IEEE Trans. Wirel. Commun. | 3 |
| 2025 | Wireless Multi-User Interactive Virtual Reality in Metaverse With Edge-Device Collaborative ComputingabstractThe immersive nature of the metaverse presents significant challenges for wireless multi-user interactive virtual reality (VR), such as ultra-low latency, high throughput and intensive computing, which place substantial demands on the wireless bandwidth and rendering resources of mobile edge computing (MEC). In this paper, we propose a wireless multi-user interactive VR with edge-device collaborative computing framework to overcome the motion-to-photon (MTP) threshold bottleneck. Specifically, we model the serial-parallel task execution in queues within a foreground and background separation architecture. The rendering indices of background tiles within the prediction window are determined, and both the foreground and selected background tiles are loaded into respective processing queues based on the rendering locations. To minimize the age of sensor information and the power consumption of mobile devices, we optimize rendering decisions and MEC resource allocation subject to the MTP constraint. To address this optimization problem, we design a safe reinforcement learning (RL) algorithm, active queue management-constrained updated projection (AQM-CUP). AQM-CUP constructs an environment suitable for queues, incorporating expired tiles actively discarded in processing buffers into its state and reward system. Experimental results demonstrate that the proposed framework significantly enhances user immersion while reducing device power consumption, and the superiority of the proposed AQM-CUP algorithm over conventional methods in terms of the training convergence and performance metrics. Caolu Xu, Zhiyong Chen 0002, Meixia Tao, Wenjun Zhang 0001 |
IEEE Trans. Wirel. Commun. | 4 |
| 2025 | Semantic-Aided Parallel Image Transmission Compatible With Practical SystemabstractIn this paper, we propose a novel semantic-aided image communication framework for supporting the compatibility with practical separation-based coding architectures. Particularly, the deep learning (DL)-based joint source-channel coding (JSCC) is integrated into the classical separate source-channel coding (SSCC) to transmit the images via the combination of semantic stream and image stream from DL networks and SSCC respectively, which we name as parallel-stream transmission. The positive coding gain stems from the sophisticated design of the JSCC encoder, which leverages the residual information neglected by the SSCC to enhance the learnable image features. Furthermore, a conditional rate adaptation mechanism is introduced to adjust the transmission rate of semantic stream according to residual, rendering the framework more flexible and efficient to bandwidth allocation. We also design a dynamic stream aggregation strategy at the receiver, which provides the composite framework with more robustness to signal-to-noise ratio (SNR) fluctuations in wireless systems compared to a single conventional codec. Finally, the proposed framework is verified to surpass the performance of both traditional and DL-based competitors in a large range of scenarios and meanwhile, maintains lightweight in terms of the transmission and computational complexity of semantic stream, which exhibits the potential to be applied in real systems. Mingkai Xu, Yongpeng Wu 0001, Yuxuan Shi 0001, Xiang-Gen Xia 0001, Mérouane Debbah, Wenjun Zhang 0001, Ping Zhang 0003 |
IEEE Trans. Wirel. Commun. | 6 |
| 2024 | RIS-Aided Receive Generalized Spatial Modulation Design with Reflecting ModulationabstractSpatial modulation (SM) transmits additional information bits by the selection of antennas. Generalized spatial modulation (GSM), as an advanced type of SM, can be divided into diversity and multiplexing (MUX) schemes according to the symbols carried on the selected antennas are identical or different. Recently, reconfigurable intelligent surface (RIS) assisted SM exhibits better reception performance compared to conventional SM. To overcome the limitations of SM, this paper combines GSM with RIS and proposes the RIS-aided receive generalized spatial modulation (RIS-RGSM) scheme. The RIS-RGSM diversity scheme is realized via a simple improvement based on the state-of-the-art scheme. To further increase the transmission rate, a novel RIS-RGSM MUX scheme is proposed, where the reflection phase shifts and on/off states of RIS elements are configured to achieve bit mapping. The theoretical bit error rate (BER) of the proposed scheme is derived and agrees well with the simulation results. Numerical simulations show that the RIS-RGSM MUX scheme has better BER performance than the diversity scheme. The proposed scheme can significantly increase the transmission rate and maintain good performance compared to the existing scheme under a limited number of antennas. Xinghao Guo, Yin Xu 0001, Hanjiang Hong, De Mi, Ruiqi Liu 0002, Dazhi He, Wenjun Zhang 0001, Yi-Yan Wu |
GLOBECOM | 7 |
| 2024 | MambaJSCC: Deep Joint Source-Channel Coding with Visual State Space ModelabstractLightweight and efficient neural network models for deep joint source-channel coding (JSCC) are crucial for semantic communications. In this paper, we design a novel JSCC scheme named MambaJSCC, which utilizes a visual state space model with channel adaptation (VSSM-CA) block as its backbone for transmitting images over wireless channels. The VSSM-CA block utilizes VSSM to integrate images with the state space, enabling feature extraction and encoding processes to operate with linear complexity. It also incorporates channel state information (CSI) via a newly proposed CSI embedding method. This method deploys a shared CSI encoding module within both the encoder and decoder to encode and inject the CSI into each VSSM-CA block, improving the adaptability of a single model to varying channel conditions. Experimental results show that MambaJSCC not only outperforms Swin Transformer based JSCC (SwinJSCC) but also significantly reduces parameter size, computational overhead, and inference delay (ID). In particular, with employing an equal number of the VSSM-CA blocks and the Swin Transformer blocks, MambaJSCC achieves a 0.48 dB gain in peak-signal-to-noise ratio (PSNR) while requiring only 53.3% multiply-accumulate operations, 53.8% of the parameters, and 44.9% of ID. Tong Wu 0003, Zhiyong Chen 0002, Meixia Tao, Xiaodong Xu 0001, Wenjun Zhang 0001, Ping Zhang 0003 |
GLOBECOM | 5 |
| 2024 | Unsourced Random Access in MIMO Quasi-Static Rayleigh Fading Channels with Finite BlocklengthabstractThis paper explores the fundamental limits of unsourced random access (URA) with a random and unknown number$\mathrm{K}_{a}$of active users in MIMO quasi-static Rayleigh fading channels. First, we derive an upper bound on the probability of incorrectly estimating the number of active users. We prove that it exponentially decays with the number of receive antennas and eventually vanishes, whereas reaches a plateau as the power and blocklength increase. Then, we derive non-asymptotic achievability and converse bounds on the minimum energy-per-bit required by each active user to reliably transmit$J$bits with blocklength$n$. Numerical results verify the tightness of our bounds, suggesting that they provide benchmarks to evaluate existing schemes. The extra required energy-per-bit due to the uncertainty of the number of active users decreases as$\mathbb{E}[\mathrm{K}_{a}]$increases. Compared to random access with individual codebooks, the URA paradigm achieves higher spectral and energy efficiency. Moreover, using codewords distributed on a sphere is shown to outperform the Gaussian random coding scheme in the non-asymptotic regime. Junyuan Gao, Yongpeng Wu 0001, Giuseppe Caire, Wei Yang 0001, Wenjun Zhang 0001 |
ISIT | 5 |
| 2024 | Pioneer: Offline Reinforcement Learning based Bandwidth Estimation for Real-Time CommunicationabstractFor Real-time Communication (RTC), Bandwidth Estimation (BWE) is crucial for enhancing user Quality of Experience (QoE) by ensuring efficient bandwidth utilization and low latency. Recent advancements have shifted towards machine learning based algorithms, particularly online reinforcement leanring (RL), to dynamically infer future bandwidth using statistical data. However, challenges such as dependency on training settings, the necessity for extensive trial and error, and instability in complex state spaces hinder their efficacy. To address these limitations, we propose Pioneer, a novel offline RL framework for BWE in RTC systems. Unlike its predecessors, Pioneer eliminates the need for real-time environment interaction during training and achieves good performance through lightweight training. Our framework consists of a Trajectory Sampler for state information preprocessing and a Bandwidth Estimator based on offline RL model. Our test results on offline datasets show that Pioneer can achieve better performance than expert algorithms. We also tested Pioneer on online simulation platforms, and Pioneer can improve QoE by 9% compared to other offline algorithm, demonstrating good robustness. Bingcong Lu, Jun Xu 0040, Rong Xie 0004, Li Song 0001, Wenjun Zhang 0001 |
MMSys | 6 |
| 2024 | AsymLLIC: Asymmetric Lightweight Learned Image CompressionabstractLearned image compression (LIC) methods often employ symmetrical encoder and decoder architectures, evitably increasing decoding time. However, practical scenarios demand an asymmetric design, where the decoder requires low complexity to cater to diverse low-end devices, while the encoder can accommodate higher complexity to improve coding performance. In this paper, we propose an asymmetric lightweight learned image compression (AsymLLIC) architecture with a novel training scheme, enabling the gradual substitution of complex decoding modules with simpler ones. Building upon this approach, we conduct a comprehensive comparison of different decoder network structures to strike a better trade-off between complexity and compression performance. Experiment results validate the efficiency of our proposed method, which not only achieves comparable performance to VVC but also offers a lightweight decoder with only 51.47 GMACs computation and 19.65M parameters. Furthermore, this design methodology can be easily applied to any LIC models, enabling the practical deployment of LIC techniques. Shen Wang 0013, Zhengxue Cheng, Donghui Feng 0003, Guo Lu, Li Song 0001, Wenjun Zhang 0001 |
VCIP | 6 |
| 2024 | Design of Capacity-Approaching Constellation and Pre-scaling for Spatial ModulationabstractSpatial Modulation (SM), as a type of index modulation (IM), can utilize the index of the transmit antenna (TA) to transmit additional information. In this paper, to improve the performance of SM, a non-uniform constellation (NUC) and pre-scaling coefficients optimization design scheme is proposed. The bit-interleaved coded modulation (BICM) capacity calculation formula of SM system is firstly derived. The constellation and pre-scaling coefficients are optimized by maximizing the BICM capacity without channel state information (CSI) feedback. Optimization results are given for the multiple-input-single-output (MISO) system with Rayleigh channel. Simulation result shows the proposed scheme provides a meaningful performance gain compared to conventional SM system without CSI feedback. The proposed optimization design scheme is a general scheme that can be used as a reference and easily extended to more scenarios with various SM schemes, so it is a promising design paradigm for future wireless communication to achieve high-efficiency. Xinghao Guo, Yin Xu 0001, Hanjiang Hong, Size Peng, Dazhi He, Wenjun Zhang 0001, Yi-Yan Wu |
VTC Spring | 6 |
| 2024 | Unsupervised Learning Based Symbol-Level Precoding Design for Amplitude Phase ModulationabstractThe symbol-level precoding (SLP) technique can enhance the performance in multi-user wireless communication systems because of its ability to convert harmful multi-user interference (MUI) into beneficial ones. However, the tremendous computational complexity of conventional symbol-level precoding designs severely hinders practical implementations. This paper proposes an SLP design scheme based on unsupervised learning in a multiple-input multiple-output (MIMO) downlink system. In the SLP design scheme, the loss function is first designed to improve performance by pushing the received signal further away from the decision boundaries into a constructive region. An efficient symbol-level precoding network (SLP-Net) is introduced to optimize the SLP under the power constraint and adapt amplitude phase modulation. Numerical results highlight that the optimized SLP design scheme provides meaningful performance gain in the MIMO downlink system. The proposed SLP design scheme can be an efficient technology to perform better in the future 6G. Liangyuan Zhao, Hao Ju 0002, XiaoWu Ou, Yin Xu 0001, Dazhi He, Sung Ik Park, Namho Hur, Wenjun Zhang 0001 |
VTC Fall | 8 |
| 2024 | Decentralization of Tomlinson-Harashima Precoding for MU-MIMO SystemabstractMulti-User Multiple-Input Multiple-Output (MU-MIMO) antenna arrays are considered a crucial technology for future wireless communication systems. However, precoding for MU-MIMO meets significant challenges. To tackle this issue, this paper introduces a novel star decentralized precoding algorithm, aiming to decentralize part of the precoded calculations from the central unit (CU) to the decentralized units (DUs) and reduce the computational complexity of the CU. Then, we apply the Zero Forcing Tomlinson-Harashima precoding (ZF-THP) algorithm to star decentralized baseband processing (DBP) for enhanced transmission rate, and this algorithm has the same performance as centralized precoding but with reduced CU complexity. Furthermore, we propose the star decentralized minimum mean square error THP (sDMMSE-THP) algorithm to enhance system performance further. Extensive simulation data validate the effectiveness of our proposed scheme. Yin Xu 0001, Guanli Yi, Dazhi He, Haoyang Li 0004, XiaoWu Ou, Yunfeng Guan 0001, Wenjun Zhang 0001 |
VTC Fall | 8 |
| 2024 | Physical layer signal processing for XR communications and systems
Yongpeng Wu 0001, Mai Xu, Guangtao Zhai, Wenjun Zhang 0001 |
Sci. China Inf. Sci. | 4 |
| 2024 | R-PMAC: A Robust Preamble-Based MAC Mechanism Applied in Industrial Internet of ThingsabstractThis article proposes a novel media access control (MAC) mechanism, called the robust preamble-based MAC mechanism (R-PMAC), which can be applied to power line communication (PLC) networks in the context of the Industrial Internet of Things (IIoT). Compared with other MAC mechanisms, such as P-MAC and the MAC layer of IEEE1901.1, R-PMAC has higher networking speed. Besides, it supports whitelist authentication and functions properly in the presence of data frame loss. First, we outline three basic mechanisms of R-PMAC, containing precise time difference calculation, preambles generation, and short ID allocation. Second, we elaborate its networking process of single layer and multiple layers. Third, we illustrate its robust mechanisms, including collision handling and data retransmission. Moreover, a low-cost hardware platform is established to measure the time of connecting hundreds of PLC nodes for the R-PMAC, P-MAC, and IEEE1901.1 mechanisms in a real power line environment. The experiment results show that R-PMAC outperforms the other mechanisms by achieving a 50% reduction in networking time. These findings indicate that the R-PMAC mechanism holds great potential for quickly and effectively building a PLC network in actual industrial scenarios. Biqian Feng, Yongpeng Wu 0001, Zhen Gao 0001, Wenjun Zhang 0001 |
IEEE Internet Things J. | 5 |
| 2024 | Massive Unsourced Random Access for Near-Field CommunicationsabstractThis paper investigates the unsourced random access (URA) problem with a massive multiple-input multiple-output receiver that serves wireless devices in the near-field of radiation. We employ an uncoupled transmission protocol without appending redundancies to the slot-wise encoded messages. To exploit the channel sparsity for block length reduction while facing the collapsed sparse structure in the angular domain of near-field channels, we propose a sparse channel sampling method that divides the angle-distance (polar) domain based on the maximum permissible coherence. Decoding starts with retrieving active codewords and channels from each slot. We address the issue by leveraging the structured channel sparsity in the spatial and polar domains and propose a novel turbo-based recovery algorithm. Furthermore, we investigate an off-grid compressed sensing method to refine discretely estimated channel parameters over the continuum that improves the detection performance. Afterward, without the assistance of redundancies, we recouple the separated messages according to the similarity of the users’ channel information and propose a modifiedK-medoids method to handle the constraints and collisions involved in channel clustering. Simulations reveal that via exploiting the channel sparsity, the proposed URA scheme achieves high spectral efficiency and surpasses existing multi-slot-based schemes. Moreover, with more measurements provided by the overcomplete channel sampling, the near-field-suited scheme outperforms its counterpart of the far-field. Xinyu Xie, Yongpeng Wu 0001, Jianping An, Derrick Wing Kwan Ng, Chengwen Xing, Wenjun Zhang 0001 |
IEEE Trans. Commun. | 6 |
| 2024 | Coarse- and Fine-Grained Fusion Hierarchical Network for Hole Filling in View SynthesisabstractDepth image-based rendering (DIBR) techniques play an essential role in free-viewpoint videos (FVVs), which generate the virtual views from a reference 2D texture video and its associated depth information. However, the background regions occluded by the foreground in the reference view will be exposed in the synthesized view, resulting in obvious irregular holes in the synthesized view. To this end, this paper proposes a novel coarse and fine-grained fusion hierarchical network (CFFHNet) for hole filling, which fills the irregular holes produced by view synthesis using the spatial contextual correlations between the visible and hole regions. CFFHNet adopts recurrent calculation to learn the spatial contextual correlation, while the hierarchical structure and attention mechanism are introduced to guide the fine-grained fusion of cross-scale contextual features. To promote texture generation while maintaining fidelity, we equip CFFHNet with a two-stage framework involving an inference sub-network to generate the coarse synthetic result and a refinement sub-network for refinement. Meanwhile, to make the learned hole-filling model better adaptable and robust to the "foreground penetration" distortion, we trained CFFHNet by generating a batch of training samples by adding irregular holes to the foreground and background connection regions of high-quality images. Extensive experiments show the superiority of our CFFHNet over the current state-of-the-art DIBR methods. The source code will be available at https://github.com/wgc-vsfm/view-synthesis-CFFHNet. Guangcheng Wang, Kui Jiang, Ke Gu 0001, Hongyan Liu 0004, Hantao Liu, Wenjun Zhang 0001 |
IEEE Trans. Image Process. | 6 |
| 2024 | Real-Time Free Viewpoint Video Synthesis System Based on DIBR and a Depth Estimation NetworkabstractDepth image-based rendering (DIBR) view synthesis is the most widely employed method in real-time FVV research. Despite recent progress, most DIBR-based FVV synthesis approaches are not sufficiently simple and effective in filling holes and artifacts. Additionally, they use RGB-D cameras, which are difficult to widely adopt or take considerable time to estimate high-quality depth images. This paper introduces a real-time FVV synthesis system based on DIBR and a depth estimation network. This system includes a 12-view synchronous camera system, a new multistage depth estimation network, a new GPU-accelerated DIBR algorithm, and a virtual view parameter generation method. This system provides the first real-time FVV solution for background-fixed fields based on DIBR and a depth estimation network. It can infer depth images for all camera views and synthesize any virtual view along the horizontal circular arc of the camera rig in real time. To our knowledge, we are the first to introduce background models and foreground masks and a refined multistage structure to address real-time high-quality depth estimation and DIBR FVV synthesis. We also build a high-quality multiview RGB-D synchronous dataset that has promising DIBR FVV synthesis performance to train and evaluate our system. The experimental results demonstrate the real-time and better performance of the proposed system. Shuai Guo 0002, Jingchuan Hu, Kai Zhou 0016, Jionghao Wang, Li Song 0001, Rong Xie 0004, Wenjun Zhang 0001 |
IEEE Trans. Multim. | 7 |
| 2024 | Joint User Association, Resource Allocation, and Beamforming in RIS-Assisted Multi-Server MEC SystemsabstractMulti-access edge computing (MEC) is a promising solution to supporting resource-intensive applications on mobile devices (MDs), which enables computation offloading from MDs to edge servers at their proximities. However, the quality of the communication links and the limited communication and computing resources significantly impact the performance of MEC systems. In this paper, we leverage the emerging reconfigurable intelligent surfaces (RISs) to assist the computation offloading and balance the computing workloads in a multi-server MEC system with limited communication and computing resources. Specifically, when a nearby edge server is overwhelmed by multiple computing tasks, some MDs can be redirected to potentially distant but lighter-loaded edge servers by employing passive beamforming enabled by RISs. Thus, to maximize the task completion rate, we formulate a joint optimization problem for user association, passive beamforming at RISs, receive beamforming at BSs, and computing resource allocation on edge servers. Since the problem is a mixed integer nonlinear programming (MINLP), which is challenging to solve, we first decompose it into two tractable subproblems through the block coordinate descent (BCD) technique and then solve them by the penalty dual decomposition (PDD) method and a swap matching-based algorithm, respectively. Numerical results demonstrate that the task completion rate can be significantly increased by incorporating RISs into multi-server MEC systems. Besides, the proposed algorithms outperform other benchmark schemes in terms of both the task completion rate and the design complexity. Wen He 0001, Dazhi He, Xianhao Chen, Yuguang Fang, Wenjun Zhang 0001 |
IEEE Trans. Wirel. Commun. | 6 |
| 2024 | CDDM: Channel Denoising Diffusion Models for Wireless Semantic CommunicationsabstractDiffusion models (DM) can gradually learn to remove noise, which have been widely used in artificial intelligence generated content (AIGC) in recent years. The property of DM for eliminating noise leads us to wonder whether DM can be applied to wireless communications to help the receiver mitigate the channel noise. To address this, we propose channel denoising diffusion models (CDDM) for semantic communications over wireless channels in this paper. CDDM can be applied as a new physical layer module after the channel equalization to learn the distribution of the channel input signal, and then utilizes this learned knowledge to remove the channel noise. We derive corresponding training and sampling algorithms of CDDM according to the forward diffusion process specially designed to adapt the channel models and theoretically prove that the well-trained CDDM can effectively reduce the conditional entropy of the received signal under small sampling steps. Moreover, we apply CDDM to a semantic communications system based on joint source-channel coding (JSCC) for image transmission and design a three-stage training algorithm for combining them. Extensive experimental results demonstrate that CDDM can further reduce the mean square error (MSE) after minimum mean square error (MMSE) equalizer, and the joint CDDM and JSCC system achieves better performance than the JSCC system, the traditional JPEG2000 with low-density parity-check (LDPC) code approach and other benchmarks in diverse scenarios. Tong Wu 0003, Zhiyong Chen 0002, Dazhi He, Liang Qian, Yin Xu 0001, Meixia Tao, Wenjun Zhang 0001 |
IEEE Trans. Wirel. Commun. | 7 |
| 2024 | Robust Image Semantic Coding With Learnable CSI Fusion Masking Over MIMO Fading ChannelsabstractThough achieving marvelous progress in various scenarios, existing semantic communication frameworks mainly consider single-input single-output Gaussian channels or Rayleigh fading channels, neglecting the widely-used multiple-input multiple-output (MIMO) channels, which hinders the application into practical systems. One common solution to combat MIMO fading is to utilize feedback MIMO channel state information (CSI). In this paper, we incorporate MIMO CSI into system designs from a new perspective and propose the learnable CSI fusion semantic communication (LCFSC) framework, where CSI is treated as side information by the semantic extractor to enhance the semantic coding. To avoid feature fusion due to abrupt combination of CSI with features, we present a non-invasive CSI fusion multi-head attention module inside the Swin Transformer. With the learned attention masking map determined by both source and channel states, more robust attention distribution could be generated. Furthermore, the percentage of mask elements could be flexibly adjusted by the learnable mask ratio, which is produced based on the conditional variational interference in an unsupervised manner. In this way, CSI-aware semantic coding is achieved through learnable CSI fusion masking. Experiment results testify the superiority of LCFSC over traditional schemes and state-of-the-art Swin Transformer-based semantic communication frameworks in MIMO fading channels. Bingyan Xie, Yongpeng Wu 0001, Yuxuan Shi 0001, Wenjun Zhang 0001, Shuguang Cui, Mérouane Debbah |
IEEE Trans. Wirel. Commun. | 4 |
| 2023 | AudioEar: Single-View Ear Reconstruction for Personalized Spatial AudioabstractSpatial audio, which focuses on immersive 3D sound rendering, is widely applied in the acoustic industry. One of the key problems of current spatial audio rendering methods is the lack of personalization based on different anatomies of individuals, which is essential to produce accurate sound source positions. In this work, we address this problem from an interdisciplinary perspective. The rendering of spatial audio is strongly correlated with the 3D shape of human bodies, particularly ears. To this end, we propose to achieve personalized spatial audio by reconstructing 3D human ears with single-view images. First, to benchmark the ear reconstruction task, we introduce AudioEar3D, a high-quality 3D ear dataset consisting of 112 point cloud ear scans with RGB images. To self-supervisedly train a reconstruction model, we further collect a 2D ear dataset composed of 2,000 images, each one with manual annotation of occlusion and 55 landmarks, named AudioEar2D. To our knowledge, both datasets have the largest scale and best quality of their kinds for public use. Further, we propose AudioEarM, a reconstruction method guided by a depth estimation network that is trained on synthetic data, with two loss functions tailored for ear data. Lastly, to fill the gap between the vision and acoustics community, we develop a pipeline to integrate the reconstructed ear mesh with an off-the-shelf 3D human body and simulate a personalized Head-Related Transfer Function (HRTF), which is the core of spatial audio rendering. Code and data are publicly available in https://github.com/seanywang0408/AudioEar. Bingbing Ni, Wenjun Zhang 0001, Jinxian Liu, Teng Li 0001 |
AAAI | 5 |
| 2023 | Boosting Point Clouds Rendering via Radiance MappingabstractRecent years we have witnessed rapid development in NeRF-based image rendering due to its high quality. However, point clouds rendering is somehow less explored. Compared to NeRF-based rendering which suffers from dense spatial sampling, point clouds rendering is naturally less computation intensive, which enables its deployment in mobile computing device. In this work, we focus on boosting the image quality of point clouds rendering with a compact model design. We first analyze the adaption of the volume rendering formulation on point clouds. Based on the analysis, we simplify the NeRF representation to a spatial mapping function which only requires single evaluation per pixel. Further, motivated by ray marching, we rectify the the noisy raw point clouds to the estimated intersection between rays and surfaces as queried coordinates, which could avoid spatial frequency collapse and neighbor point disturbance. Composed of rasterization, spatial mapping and the refinement stages, our method achieves the state-of-the-art performance on point clouds rendering, outperforming prior works by notable margins, with a smaller model size. We obtain a PSNR of 31.74 on NeRF-Synthetic, 25.88 on ScanNet and 30.81 on DTU. Code and data are publicly available in https://github.com/seanywang0408/RadianceMapping. Bingbing Ni, Teng Li 0001, Kai Chen 0006, Wenjun Zhang 0001 |
AAAI | 6 |
| 2023 | Freestyle Layout-to-Image SynthesisabstractTypical layout-to-image synthesis (LIS) models generate images for a closed set of semantic classes, e.g., 182 common objects in COCO-Stuff. In this work, we explore the freestyle capability of the model, i.e., how far can it generate unseen semantics (e.g., classes, attributes, and styles) onto a given layout, and call the task Freestyle LIS (FLIS). Thanks to the development of large-scale pre-trained language-image models, a number of discriminative models (e.g., image classification and object detection) trained on limited base classes are empowered with the ability of unseen class prediction. Inspired by this, we opt to leverage large-scale pre-trained text-to-image diffusion models to achieve the generation of unseen semantics. The key challenge of FLIS is how to enable the diffusion model to synthesize images from a specific layout which very likely violates its pre-learned knowledge, e.g., the model never sees “a unicorn sitting on a bench” during its pre-training. To this end, we introduce a new module called Rectified Cross-Attention (RCA) that can be conveniently plugged in the diffusion model to integrate semantic masks. This “plug-in” is applied in each cross-attention layer of the model to rectify the attention maps between image and text tokens. The key idea of RCA is to enforce each text token to act on the pixels in a specified region, allowing us to freely put a wide variety of semantics from pre-trained knowledge (which is general) onto the given layout (which is specific). Extensive experiments show that the proposed diffusion network produces realistic and freestyle layout-to-image generation results with diverse text inputs, which has a high potential to spawn a bunch of interesting applications. Code is available at https://github.com/essunny310/FreestyleNet. Zhiwu Huang, Qianru Sun, Li Song 0001, Wenjun Zhang 0001 |
CVPR | 5 |
| 2023 | Frequency-Modulated Point Cloud Rendering with Easy EditingabstractWe develop an effective point cloud rendering pipeline for novel view synthesis, which enables high fidelity local detail reconstruction, real-time rendering and user-friendly editing. In the heart of our pipeline is an adaptive frequency modulation module called Adaptive Frequency Net (AFNet), which utilizes a hypernetwork to learn the local texture frequency encoding that is consecutively injected into adaptive frequency activation layers to modulate the implicit radiance signal. This mechanism improves the frequency expressive ability of the network with richer frequency basis support, only at a small computational budget. To further boost performance, a preprocessing module is also proposed for point cloud geometry optimization via point opacity estimation. In contrast to implicit rendering, our pipeline supports high-fidelity interactive editing based on point cloud manipulation. Extensive experimental results on NeRF-Synthetic, ScanNet, DTU and Tanks and Temples datasets demonstrate the superior performances achieved by our method in terms of PSNR, SSIM and LPIPS, in comparison to the state-of-the-art. Code is released at https://github.com/yizhangphd/FreqPCR. Bingbing Ni, Wenjun Zhang 0001, Teng Li 0001 |
CVPR | 4 |
| 2023 | Boosting Video Object Segmentation via Space-Time Correspondence LearningabstractCurrent top-leading solutions for video object segmentation (VOS) typically follow a matching-based regime: for each query frame, the segmentation mask is inferred according to its correspondence to previously processed and the first annotated frames. They simply exploit the supervisory signals from the groundtruth masks for learning mask prediction only, without posing any constraint on the space-time correspondence matching, which, however, is the fundamental building block of such regime. To alleviate this crucial yet commonly ignored issue, we devise a correspondence-aware training framework, which boosts matching-based VOS solutions by explicitly encouraging robust correspondence matching during network learning. Through comprehensively exploring the intrinsic coherence in videos on pixel and object levels, our algorithm reinforces the standard, fully supervised training of mask segmentation with label-free, contrastive correspondence learning. Without neither requiring extra annotation cost during training, nor causing speed delay during deployment, nor incurring architectural modification, our algorithm provides solid performance gains on four widely used benchmarks, i.e., DAVIS2016&2017, and YouTube-VOS2018&2019, on the top of famous matching-based VOS solutions. Liulei Li, Wenguan Wang, Rong Xie 0004, Li Song 0001, Wenjun Zhang 0001 |
CVPR | 6 |
| 2023 | A Graph-Based Collision Resolution Scheme for Asynchronous Unsourced Random AccessabstractThis paper investigates the multiple-input-multiple-output (MIMO) massive unsourced random access in an asynchronous orthogonal frequency division multiplexing (OFDM) system, with both timing and frequency offsets (TFO) and non-negligible user collisions. The proposed coding framework splits the data into two parts encoded by sparse regression code (SPARC) and low-density parity check (LDPC) code. Multistage orthogonal pilots are transmitted in the first part to reduce collision density. Unlike existing schemes requiring a quantization codebook with a large size for estimating TFO, we establish a graph-based channel reconstruction and collision resolution (GB-CR2) algorithm to iteratively reconstruct channels, resolve collisions, and compensate for TFO rotations on the formulated graph jointly among multiple stages. We further propose to leverage the geometric characteristics of signal constellations to correct TFO estimations. Exhaustive simulations demonstrate remarkable performance superiority in channel estimation and data recovery with substantial complexity reduction compared to state-of-the-art schemes. Tianya Li, Yongpeng Wu 0001, Wenjun Zhang 0001, Xiang-Gen Xia 0001, Chengshan Xiao |
GLOBECOM | 3 |
| 2023 | Joint Device Identification, Channel Estimation, and Signal Detection for LEO Satellite-Enabled Random AccessabstractThis paper investigates joint device identification, channel estimation, and signal detection for LEO satellite-enabled grant-free random access, where a multiple-input multiple-output (MIMO) system with orthogonal time-frequency space modulation (OTFS) is utilized to combat the dynamics of the terrestrial-satellite link (TSL). We divide the receiver structure into three modules: first, a linear module for identifying active devices, which leverages the generalized approximate message passing (GAMP) algorithm to eliminate inter-user interference in the delay-Doppler domain; second, a non-linear module adopting the message passing algorithm to jointly estimate channel and detect transmit signals; the third aided by Markov random field (MRF) aims to explore the three dimensional block sparsity of channel in the delay-Doppler-angle domain. The soft information is exchanged iteratively between these three modules by careful scheduling. Furthermore, the expectation-maximization algorithm is embedded to learn the hyperparameters in prior distributions. Simulation results demonstrate that the proposed scheme outperforms the conventional methods significantly in terms of activity error rate, channel estimation accuracy, and symbol error rate. Boxiao Shen, Yongpeng Wu 0001, Wenjun Zhang 0001, Symeon Chatzinotas, Björn Ottersten 0001 |
GLOBECOM | 3 |
| 2023 | CDDM: Channel Denoising Diffusion Models for Wireless CommunicationsabstractDiffusion models (DM) can gradually learn to re-move noise, which have been widely used in artificial intelligence generated content (AIGC) in recent years. The property of DM for removing noise leads us to wonder whether DM can be applied to wireless communications to help the receiver eliminate the channel noise. To address this, we propose channel denoising diffusion models (CDDM) for wireless communications in this paper. CDDM can be applied as a new physical layer module after the channel equalization to learn the distribution of the channel input signal, and then utilizes this learned knowledge to remove the channel noise. We design corresponding training and sampling algorithms for the forward diffusion process and the reverse sampling process of CDDM. Moreover, we apply CDDM to a semantic communications system based on joint source-channel coding (JSCC). Experimental results demonstrate that CDDM can further reduce the mean square error (MSE) after minimum mean square error (MMSE) equalizer, and the joint CDDM and JSCC system achieves better performance than the JSCC system and the traditional JPEG2000 with low-density parity-check (LDPC) code approach. Tong Wu 0003, Zhiyong Chen 0002, Dazhi He, Liang Qian, Yin Xu 0001, Meixia Tao, Wenjun Zhang 0001 |
GLOBECOM | 7 |
| 2023 | Fusion-Based Multi-User Semantic Communications for Wireless Image Transmission Over Degraded Broadcast ChannelsabstractDegraded broadcast channels (DBC) are a typical multiuser communication scenario. There exist classic transmission methods, such as superposition coding with successive interference cancellation, to achieve the DBC capacity region. However, semantic communication method over DBC remains lack of in-depth research. To address this, we design a semantic communications system for wireless image transmission over DBC in this paper. The proposed architecture supports a transmitter extracting semantic features for two users separately, and learns to dynamically fuse these semantic features into a joint latent representation for broadcasting. The key here is to design a flexible image semantic fusion (FISF) module to fuse the semantic features of two users, and to use a multi-layer perceptron (MLP) based neural network to adjust the weights of different user semantic features for flexible adaptability to different users channels. Experiments present the semantic performance region based on the peak signal-to-noise ratio (PSNR) of both users, and show that the proposed system dominates the traditional methods. Tong Wu 0003, Zhiyong Chen 0002, Meixia Tao, Bin Xia 0001, Wenjun Zhang 0001 |
GLOBECOM | 5 |
| 2023 | Edge-Device Collaborative Rendering for Wireless Multi-User Interactive Virtual Reality in MetaverseabstractThe immersive nature of the metaverse poses higher requirements on virtual reality (VR). In multi-user interactive VR, achieving real-time rendering of heterogeneous foreground objects with ultra-low latency is a challenge. In this paper, we propose an edge-device collaborative rendering framework based on the real-time computer graphics (CG) execution. We design a multi-object real-time rendering workflow and obtain the corresponding motion-to-photon (MTP) latency. We joint optimize the rendering positions of foreground objects and the bandwidth resource to minimize the MTP latency. A smooth maximum joint rendering scheme is proposed to address the non-differentiable objective function and convert the non-convex problem into convex subproblems. Numerical results illustrate that the proposed scheme significantly improves the frame rate by compensating the bottleneck term in the MTP latency. Caolu Xu, Zhiyong Chen 0002, Meixia Tao, Wenjun Zhang 0001 |
GLOBECOM | 4 |
| 2023 | Dual-Head Fusion Network for Image EnhancementabstractImage enhancement algorithms have made great progress recently. However, most existing methods tend to construct a uniform enhancer for the color transformation of all pixels and ignore the local context information which is significant for photographs, causing unsatisfactory results. To solve these issues, we propose a novel dual-head fusion network for image enhancement, which synthetically considers both global scenario and local content information. Our network consists of four lightweight modules. We first develop a dual-head feature extraction module to extract the global condition vector and spatial context map. After that, we propose a context-aware retouching module and a global color rendering module to generate latent results. Finally, we employ the spatial attention based fusion module to adaptively aggregate the latent results. Experiments on public datasets show that our method consistently achieves the best results compared with SOTA methods both quantitatively and qualitatively. Hengsheng Zhang, Li Song 0001, Rong Xie 0004, Wenjun Zhang 0001 |
ICASSP | 5 |
| 2023 | Iterative Channel Estimation for OTFS Using ZC Sequence with Low Peak-to-Average Power RatioabstractFor low earth orbit (LEO) satellite communications, the robustness to high mobility and the low peak-to-average power ratio (PAPR) are two essential requirements. Orthogonal time frequency space (OTFS) modulation outperforms orthogonal frequency division multiplexing (OFDM) in high-mobility scenarios with doubly selective channels. However, the existing channel estimation schemes for OTFS usually contain high-power pilots, which cause high PAPR, while the current PAPR reduction schemes for OTFS usually ignore the impacts on channel estimation. This work proposes a novel channel estimation scheme utilizing a Zadoff-Chu (ZC) sequence as the pilot, with which the PAPR can be effectively reduced by proper pilot alignment. Then, an iterative ZC-sequence-based estimation algorithm is proposed, which adopts a correlation-based algorithm and a message passing (MP) algorithm to detect the pilot and data alternately. Simulation results show that the proposed scheme has a significantly lower PAPR in the time domain and superior channel estimation performance. Tianyao Ma, Yin Xu 0001, XiaoWu Ou, Dazhi He, Wenjun Zhang 0001 |
ICC | 6 |
| 2023 | Learning Shape Primitives via Implicit Convexity RegularizationabstractShape primitives decomposition has been an important and long-standing task in 3D shape analysis. Prior arts heavily rely on 3D point clouds or voxel data for shape primitives extraction, which are less practical in real-world scenarios. This paper proposes to learn shape primitives from multi-view images by introducing implicit surface rendering. It is challenging since implicit shapes have a high degree of freedom, which violates the simplicity property of shape primitives. In this work, a novel regularization term named Implicit Convexity Regularization (ICR) imposed on implicit primitive learning is proposed to tackle this problem. We start with the convexity definition of general 3D shapes, and then derive the equivalent expression for implicit shapes represented by signed distance functions (SDFs). Further, instead of directly constraining the output SDF values which cause unstable optimization, we alternatively impose constraint on second order directional derivatives on line segments inside the shapes, which proves to be a tighter condition for 3D convexity. Implicit primitives constrained by the proposed ICR are combined into a whole object via softmax-weighted-sum operation over all primitive SDFs. Experiments on synthetic and real-world datasets show that our method is able to decompose objects into simple and reasonable shape primitives without the need of segmentation labels or 3D data. Code and data is publicly available in https://github.com/seanywang0408/ICR. Kai Chen 0006, Teng Li 0001, Wenjun Zhang 0001, Bingbing Ni |
ICCV | 5 |
| 2023 | A Novel Labeling Scheme for Neural Belief Propagation in Polar CodesabstractRecently, deep learning has been adopted to improve the performance of the belief propagation algorithm in polar codes, namely, the neural belief propagation decoder. By attaching trainable parameters to each edge on a tanner graph, this decoder effectively weakens the effect of short cycles. In this paper, we will show that the decoder can be further improved with our newly-designed labeling scheme. Instead of groundtruth transmitted codewords, our scheme utilizes codewords generated from a high-performance decoder (i.e., noise-aided belief propagation list decoder) to construct loss functions. Besides, to reduce the labeling complexity, the ground-truth transmitted codewords are integrated into the decoder as a list and compared with the other lists with the principle of minimum Euclidean distance. Numerical simulation results demonstrate that the proposed labeling scheme can improve the decoding performance by up to 0.2 dB with the same run time complexity and model size. Hao Ju 0002, Yin Xu 0001, Dazhi He, Wenjun Zhang 0001 |
IWCMC | 5 |
| 2023 | A Deep Learning based Multi-edge-type decoding algorithm for 5G NR LDPC codesabstractLow-density parity-check(LDPC) code has been selected as the channel coding method by 5G NR because of its excellent error-correcting performance. To further improve the performance of LDPC decoding, this paper proposes a neural normalized min-sum(NNMS) algorithm based on multi-edge-type(MET). Based on the LLR convergence analysis of the protograph matrix of 5G NR, the base matrix is divided into several independent regions. Each part is assigned a unique scaling factor at different iterations. To verify the effectiveness of the proposed algorithm, We use two parity-check matrixes(PCM) derived from different base graphs in simulations. The results show that the proposed algorithm performs at most 0.45dB better than BP, 0.37dB better than NMS, and 0.25dB better than OMS, respectively, when the frame error rate (FER) is at 1$0^{-5}$ level over additive white Gaussian noise (AWGN) channels using BPSK modulation. Tianyu Du, Hao Ju 0002, Yin Xu 0001, Dazhi He, Wenjun Zhang 0001 |
IWCMC | 5 |
| 2023 | NeRF-SDP: Efficient Generalizable Neural Radiance Field with Scene Depth PerceptionabstractIn recent years, neural radiance fields have exhibited impressive performance in novel view synthesis. However, exploiting complex network structures to achieve generalizable NeRF usually results in inefficient rendering. Existing methods for accelerating rendering directly employ simpler inference networks or fewer sampling points, leading to unsatisfactory synthesis quality. To address the challenge of balancing rendering speed and quality in generalizable NeRF, we propose a novel framework, NeRF-SDP, which achieves both efficiency and high fidelity by introducing scene depth perception. We incorporate more scene information into the radiance field by using our proposed geometry feature extraction and depth-encoded ray transformer to improve the model’s inference capabilities with sparse points. With the aid of scene depth perception, NeRF-SDP can better understand the scene’s structure, thus better reconstructing the objects’ edges with significantly fewer artifacts. Experimental results demonstrate that NeRF-SDP achieves comparable synthesis quality to state-of-the-art methods while significantly improving rendering efficiency. Furthermore, ablation studies confirm that the depth-encoded ray transformer enhances the model’s robustness to varying numbers of sampling points. Qiuwen Wang, Shuai Guo 0002, Haoning Wu 0002, Rong Xie 0004, Li Song 0001, Wenjun Zhang 0001 |
MMAsia | 6 |
| 2023 | Communication-Efficient Framework for Distributed Image Semantic Wireless TransmissionabstractMultinode communication, which refers to the interaction among multiple devices, has attracted lots of attention in many Internet of Things (IoT) scenarios. However, its huge amounts of data flows and inflexibility for task extension have triggered the urgent requirement of communication-efficient distributed data transmission frameworks. In this article, inspired by the great superiorities on bandwidth reduction and task adaptation of semantic communications, we propose a federated learning (FL)-based semantic communication (FLSC) framework for multitask distributed image transmission with IoT devices. FL enables the design of independent semantic communication link of each user while further improves the semantic extraction and task performance through global aggregation. Each link in FLSC is composed of a hierarchical vision transformer (HVT)-based extractor and a task-adaptive translator for coarse-to-fine semantic extraction and meaning translation according to specific tasks. In order to extend the FLSC into more realistic conditions, we design a channel state information-based multiple-input–multiple-output transmission module to combat channel fading and noise. Simulation results show that the coarse semantic information can deal with a range of image-level tasks. Moreover, especially in low signal-to-noise ratio (SNR) and channel bandwidth ratio regimes, FLSC evidently outperforms the traditional scheme, e.g., about 10 peak SNR gain in the 3-dB channel condition. Bingyan Xie, Yongpeng Wu 0001, Yuxuan Shi 0001, Derrick Wing Kwan Ng, Wenjun Zhang 0001 |
IEEE Internet Things J. | 5 |
| 2023 | Deep Online Video Stabilization Using IMU SensorsabstractIn this paper, we propose a deep learning based sensor-driven method for online video stabilization. This method utilizes the Euler angles and acceleration values estimated from the gyroscope and accelerator to assist stable video reconstruction. We introduce two simple sub-networks for trajectory optimization. The first network exploits real unstable trajectories and camera acceleration values to detect shooting scenarios. This network also generates an attention mask to adaptively choose scenario-specific features. Then the second network predicts smooth camera paths based on real unstable trajectories using long short-term memory (LSTM) under the supervision of the above mask. The output of the trajectory optimization network is filtered with a two-step modification process to guarantee smoothness. The real and smoothed camera paths are then utilized as guidance to generate stable frames in a projective manner. We also capture videos with sensor data covering seven typical shooting scenarios and design a ground truth generation method to construct pseud-labels. Moreover, the trajectory smoothing network allows the use of 3- or 10-frame buffers as future information to construct a lookahead filter. Experimental results show that our online method could outperform other state-of-the-art offline methods in several shaky video clips with fewer buffer frames for both general and low-quality videos. Furthermore, our method could effectively reduce running times without performing image content analysis, and the stabilization efficiency reaches 25 fps on 1080p videos. Chen Li 0021, Li Song 0001, Rong Xie 0004, Wenjun Zhang 0001 |
IEEE Trans. Multim. | 5 |
| 2023 | Residual Quantization for Low Bit-Width Neural NetworksabstractNeural network quantization has shown to be an effective way for network compression and acceleration. However, existing binary or ternary quantization methods suffer from two major issues. First, low bit-width input/activation quantization easily results in severe prediction accuracy degradation. Second, network training and quantization are always treated as two non-related tasks, leading to accumulated parameter training error and quantization error. In this work, we introduce a novel scheme, namedResidual Quantization, to train a neural network with both weights and inputs constrained to low bit-width, e.g., binary or ternary values. On one hand, by recursively performing residual quantization, the resulting binary/ternary network is guaranteed to approximate the full-precision network with much smaller errors. On the other hand, we mathematically re-formulate the network training scheme in anEM-likemanner, which iteratively performs network quantization and parameter optimization. Duringexpectation, the low bit-width network is encouraged to approximate the full-precision network. Duringmaximization, the low bit-width network is further tuned to gain better representation capability. Extensive experiments well demonstrate that the proposed quantization scheme outperforms previous low bit-width methods and achieves much closer performance to the full-precision counterpart. Zefan Li, Bingbing Ni, Xiaokang Yang 0001, Wenjun Zhang 0001, Wen Gao 0001 |
IEEE Trans. Multim. | 4 |
| 2023 | Local Bidirection Recurrent Network for Efficient Video Deblurring with the Fused Temporal Merge ModuleabstractVideo deblurring methods exploit the correlation between consecutive blurry inputs to generate sharp frames. However, designing an effective and efficient method is a challenging problem for video deblurring. To guarantee the effectiveness and further improve the deblurring performance, we adopt the recurrent-based method as the baseline and reconsider the recurrent mechanism as well as the temporal feature alignment in the state-of-the-art methods. For the recurrent mechanism, we add the local backward connection to the global forward recurrent backbone to effectively exploit accurate future information. For the temporal alignment, we adopt a fused temporal merge module that exploits the superiority of flow-based and kernel-based methods with progressive correlation volumes estimation. In addition, we evaluate our method with both synthetic datasets (GoPro, DVD) and a realistic dataset (BSD). The experimental results demonstrate that our method achieves significant performance improvement with a slight computational cost increase against the state-of-the-art video deblurring methods. The extended ablation studies verify the effectiveness of our model. Chen Li 0021, Li Song 0001, Rong Xie 0004, Wenjun Zhang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | High-Fidelity Face Reenactment Via Identity-Matched Correspondence LearningabstractFace reenactment aims to generate an animation of a source face using the poses and expressions from a target face. Although recent methods have made remarkable progress by exploiting generative adversarial networks, they are limited in generating high-fidelity and identity-preserving results due to the inappropriate driving information and insufficiently effective animating strategies. In this work, we propose a novel face reenactment framework that achieves both high-fidelity generation and identity preservation. Instead of sparse face representations (e.g., facial landmarks and keypoints), we utilize the Projected Normalized Coordinate Code (PNCC) to better preserve facial details. We propose to reconstruct the PNCC with the source identity parameters and the target pose and expression parameters estimated by 3D face reconstruction to factor out the target identity. By adopting the reconstructed representation as the driving information, we address the problem of identity mismatch. To effectively utilize the driving information, we establish the correspondence between the reconstructed representation and the source representation based on the features extracted by an encoder network. This identity-matched correspondence is then utilized to animate the source face using a novel feature transformation strategy. The generator network is further enhanced by the proposed geometry-aware skip connection. Once trained, our model can be applied to previously unseen faces without further training or fine-tuning. Through extensive experiments, we demonstrate the effectiveness of our method in face reenactment and show that our model outperforms state-of-the-art approaches both qualitatively and quantitatively. Additionally, the proposed PNCC reconstruction module can be easily inserted into other methods and improve their performance in cross-identity face reenactment. Jun Ling, Anni Tang, Li Song 0001, Rong Xie 0004, Wenjun Zhang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2023 | Toward Visual Behavior and Attention Understanding for Augmented 360 Degree VideosabstractAugmented reality (AR) overlays digital content onto reality. In an AR system, correct and precise estimations of user visual fixations and head movements can enhance the quality of experience by allocating more computational resources for analyzing, rendering, and 3D registration on the areas of interest. However, there is inadequate research to help in understanding the visual explorations of the users when using an AR system or modeling AR visual attention. To bridge the gap between the saliency prediction on real-world scenes and on scenes augmented by virtual information, we construct the ARVR saliency dataset. The virtual reality (VR) technique is employed to simulate the real-world. Annotations of object recognition and tracking as augmented contents are blended into omnidirectional videos. The saliency annotations of head and eye movements for both original and augmented videos are collected and together constitute the ARVR dataset. We also design a model that is capable of solving the saliency prediction problem in AR. Local block images are extracted to simulate the viewport and offset the projection distortion. Conspicuous visual cues in the local block images are extracted to constitute the spatial features. The optical flow information is estimated as an important temporal feature. We also consider the interplay between virtual information and reality. The composition of the augmentation information is distinguished, and the joint effects of adversarial augmentation and complementary augmentation are estimated. The Markov chain is constructed with block images as graph nodes. In the determination of the edge weights, both the characteristics of the viewing behaviors and the visual saliency mechanisms are considered. The order of importance for block images is estimated through the state of equilibrium of the Markov chain. Extensive experiments are conducted to demonstrate the effectiveness of the proposed method. Yucheng Zhu, Xiongkuo Min, Dandan Zhu 0001, Guangtao Zhai, Xiaokang Yang 0001, Wenjun Zhang 0001, Ke Gu 0001, Jiantao Zhou 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2023 | Detecting Abrupt Change in Channel Covariance Matrix for MIMO CommunicationabstractThe acquisition of the channel covariance matrix is of paramount importance to many strategies in multiple-input-multiple-output (MIMO) communications, such as the minimum mean-square error (MMSE) channel estimation. Therefore, plenty of efficient channel covariance matrix estimation schemes have been proposed in the literature. However, an abrupt change in the channel covariance matrix may happen occasionally in practice due to the change in the scattering environment and the user location. Our paper aims to adopt the classic change detection theory to detect the change in the channel covariance matrix as accurately and quickly as possible such that the new covariance matrix can be re-estimated in time. Specifically, this paper first considers the technique of on-line change detection (also known as quickest/sequential change detection), where we need to detect whether a change in the channel covariance matrix occurs at each channel coherence time interval. Next, because the complexity of detecting the change in a high-dimension covariance matrix at each coherence time interval is too high, we devise a low-complexity off-line strategy in massive MIMO systems, where change detection is merely performed at the last channel coherence time interval of a given time period. Numerical results show that our proposed on-line and off-line schemes can detect the channel covariance change with a small delay and a low false alarm rate. Therefore, our paper theoretically and numerically verifies the feasibility of detecting the channel covariance change accurately and quickly in practice. Runnan Liu, Liang Liu 0003, Dazhi He, Wenjun Zhang 0001, Erik G. Larsson |
IEEE Trans. Wirel. Commun. | 4 |
| 2023 | Excess Distortion Exponent Analysis for Semantic-Aware MIMO Communication SystemsabstractIn this paper, the analysis of excess distortion exponent for joint source-channel coding (JSCC) in semantic-aware communication systems is presented. By introducing an unobservable semantic source, we extend the classical results by Csiszar to semantic-aware communication systems. Both upper and lower bounds of the exponent for the discrete memoryless source-channel pair are established. Moreover, an extended achievable bound of the excess distortion exponent for MIMO systems is derived. Further analysis explores how the block fading and numbers of antennas influence the exponent of semantic-aware MIMO systems. Our results offer some theoretical bounds of error decay performance and can be used to guide future semantic communications with joint source-channel coding scheme. Yuxuan Shi 0001, Shuo Shao 0001, Yongpeng Wu 0001, Wenjun Zhang 0001, Xiang-Gen Xia 0001, Chengshan Xiao |
IEEE Trans. Wirel. Commun. | 4 |
| 2022 | Remember Intentions: Retrospective-Memory-based Trajectory PredictionabstractTo realize trajectory prediction, most previous methods adopt the parameter-based approach, which encodes all the seen past-future instance pairs into model parameters. However, in this way, the model parameters come from all seen instances, which means a huge amount of irrelevant seen instances might also involve in predicting the current situation, disturbing the performance. To provide a more explicit link between the current situation and the seen instances, we imitate the mechanism of retrospective memory in neuropsychology and propose MemoNet, an instance-based approach that predicts the movement intentions of agents by looking for similar scenarios in the training data. In MemoNet, we design a pair of memory banks to explicitly store representative instances in the training set, acting as prefrontal cortex in the neural system, and a trainable memory addresser to adaptively search a current situation with similar instances in the memory bank, acting like basal ganglia. During prediction, MemoNet recalls previous memory by using the memory addresser to index related instances in the memory bank. We further propose a two-step trajectory prediction system, where the first step is to leverage MemoNet to predict the destination and the second step is to fulfill the whole trajectory according to the predicted destinations. Experiments show that the proposed MemoNet improves the FDE by 20.3%/10.2%/28.3%from the previous best method on SDD/ETH-UCY/NBA datasets. Experiments also show that our MemoNet has the ability to trace back to specific instances during prediction, promoting more interpretability. Chenxin Xu, Weibo Mao, Wenjun Zhang 0001, Siheng Chen |
CVPR | 3 |
| 2022 | Latency-Aware Collaborative Perception
Zixing Lei, Shunli Ren, Yue Hu 0011, Wenjun Zhang 0001, Siheng Chen |
ECCV (32) | 4 |
| 2022 | Linear MIMO Precoders Design for Finite Alphabet Inputs via Model-Free TrainingabstractThis paper investigates a novel method for designing linear precoders with finite alphabet inputs based on autoencoders (AE) without the knowledge of the channel model. By model-free training of the autoencoder in a multiple-input multiple-output (MIMO) system, the proposed method can effectively solve the optimization problem to design the precoders that maximize the mutual information between the channel inputs and outputs, when only the input-output information of the channel can be observed. Specifically, the proposed method regards the receiver and the precoder as two independent parameterized functions in the AE and alternately trains them using the exact and approximated gradient, respectively. Compared with previous precoders design methods, it alleviates the limitation of requiring the explicit channel model to be known. Simulation results show that the proposed method works as well as those methods under known channel models in terms of maximizing the mutual information and reducing the bit error rate. Biqian Feng, Yongpeng Wu 0001, Derrick Wing Kwan Ng, Wenjun Zhang 0001 |
GLOBECOM | 5 |
| 2022 | LEO Satellite-Enabled Grant-Free Random Access with MIMO-OTFSabstractThis paper investigates joint channel estimation and device activity detection in the LEO satellite-enabled grant-free random access systems with large differential delay and Doppler shift. In addition, the multiple-input multiple-output (MIMO) with orthogonal time-frequency space modulation (OTFS) is utilized to combat the dynamics of the terrestrial-satellite link. To simplify the computation process, we estimate the channel tensor in parallel along the delay dimension. Then, the deep learning and expectation-maximization approach are integrated into the generalized approximate message passing with cross-correlation-based Gaussian prior to capture the channel sparsity in the delay-Doppler-angle domain and learn the hyperparameters. Finally, active devices are detected by computing energy of the estimated channel. Simulation results demonstrate that the proposed algorithms outperform conventional methods. Boxiao Shen, Yongpeng Wu 0001, Wenjun Zhang 0001, Geoffrey Ye Li, Jianping An, Chengwen Xing |
GLOBECOM | 3 |
| 2022 | Low-Complexity Multi-Model CNN in-Loop Filter for AVS3abstractConvolutional neural network (CNN) has demonstrated powerful capabilities in many image/video processing tasks. In this paper, a low-complexity multi-model CNN in-loop filtering scheme is proposed for AVS3. Firstly, we carefully choose simplified ResNet as the lightweight single model of our proposed network. Subsequently, based on the selected single model, the multi-model iterative training framework is proposed to train a multi-model filter, where the network depth and the number of multi-models are customized for different ranges of bit rate to achieve the trade-off between model performance and computational complexity. Experimental results show that our method achieves on average 6.06% BD-rate reduction on Y component under all intra configuration. Compared to other CNN filters with comparable performance, our proposed multi-model filter can significantly reduce the decoder complexity, and the experimental results indicate that the decoding time can be saved by 26.6% on average. Shen Wang 0013, Yibing Fu, Li Song 0001, Wenjun Zhang 0001 |
ICASSP | 5 |
| 2022 | Representation-Agnostic Shape Fields
Jiancheng Yang, Linguo Li, Teng Li 0001, Bingbing Ni, Wenjun Zhang 0001 |
ICLR | 8 |
| 2022 | An Attention Based CNN with Temporal Hierarchical Deployment for AVS3 Inter In-loop FilteringabstractConvolutional Neural Network (CNN) based in-loop filter in video coding has demonstrated its superiority in benefiting coding efficiency and enhancing visual quality. In this paper, we develop a lightweight CNN-based in-loop filter for AVS3 encoder. The proposed network consists of several residual blocks with two attention branches, namely Dual Attention Network (DAN). The added channel attention branch and spatial attention branch can take advantage of the correlation between channels and pixels, improving the quality of reconstructed frames. In addition, by analyzing the inter prediction reference structure, we propose a temporal hierarchical deployment strategy to incorporate DAN into AVS3 video encoder. Therefore reconstructed frames with different distortions and referenced levels can be enhanced according to their temporal layer. Experiments prove the effectiveness of our strategy and results show our method achieves up to 6.57% and on average 3.64% BD-rate reduction on Y component under Random Access configuration. Yibing Fu, Shen Wang 0013, Li Song 0001, Wenjun Zhang 0001 |
ISCAS | 5 |
| 2022 | Power Allocation for LDM-based Hybrid Multicast TransmissionabstractThe development of non-orthogonal multiplexing (NOM) techniques such as layered-division-multiplexing (LDM) brings opportunities to the evolution of the fifth generation mobile networks (5G). In order to improve the performance of 5G multicast and beyond, this paper proposes an LDM-based hybrid multicast system to take advantage of both Multicast/Broadcast Single Frequency Network (MBSFN) and Single Cell Point to Multipoint (SC-PTM). Specifically, each cell in the network transmits the upper layer LDM signal in the MBSFN mode, while transmits the lower layer LDM signal in the SC-PTM mode. With the aim of maximizing system throughput, we formulate an optimization problem to allocate the transmit power between these LDM signal layers in each cell. Then, to solve the formulated problem, an algorithm based on the concave-convex procedure is proposed, which guarantees the convergence. Numerical simulation evaluates the performance of the proposed algorithm and LDM-based hybrid multicast, and demonstrates the improvement in system throughput and stream data rate. Yiwei Zhang 0015, Yin Xu 0001, Lidie Liu, Dazhi He, Wenjun Zhang 0001 |
IWCMC | 6 |
| 2022 | Design of Non-uniform Constellations in the Channel with Phase NoiseabstractThe performance of sub-TeraHertz(sub-THz) system is severely degraded by strong oscillator phase noise. High-order constellation is one of the methods in which transmission rates can be increased and non-uniform constellation (NUC) is considered to be an effective way of increasing block error rate (BLER) performance. In this paper, we design a series of constellations for single-carrier and multi-carrier systems under phase noise (PN). In order to improve the BLER performance under PN channel, we formulate an optimization problem to maximize the channel capacity by changing the complex plane coordinates of the constellation points. To calculate the channel capacity, single-carrier PN and multi-carrier PN are modeled separately in this paper. In particular, a series of derivations are carried out for multi-carrier PN. Numerical results demonstrate that NUC can be considered as a commitment technique for channels with PN, and a series of NUCs with different code rates are obtained, with a maximum gain of 4.10 dB in the single-carrier system and 1.65 dB in the multi-carrier system. Peiyi Zhao, Yin Xu 0001, Dazhi He, Hanjiang Hong, Wenjun Zhang 0001 |
IWCMC | 6 |
| 2022 | Multi-Scale Coarse-to-Fine Transformer for Frame InterpolationabstractThe majority of prevailing video interpolation methods compute flows to estimate the intermediate motion. However, accurate estimation of the intermediate motion is difficult with low-order motion model hypothesis, which induces enormous difficulties for subsequent processing. To alleviate the limitation, we propose a two-stage flow-free video interpolation architecture. Rather than utilizing pre-defined motion models, our method represents complex motion through data-driven learning. In the first stage, we analyze spatial-temporal information and generate coarse anchor frame features. In the second stage, we employ transformers to transfer neighboring features to the intermediate time steps and enhance the spatial textures. To improve the quality of coarse anchor frame features and the robustness in dealing with the multi-scale textures with large-scale motion, we propose a multi-scale architecture and transformers with variable token sizes to progressively enhance the features. The experimental results demonstrate that our model outperforms state-of-the-art methods for both single frame and multi frames interpolation tasks, and the extended ablation studies verify the effectiveness of our model. Chen Li 0021, Li Song 0001, Xueyi Zou, Jiaming Guo, Youliang Yan, Wenjun Zhang 0001 |
ACM Multimedia | 6 |
| 2022 | OCR-Pose: Occlusion-aware Contrastive Representation for Unsupervised 3D Human Pose EstimationabstractOcclusion is a significant problem in 3D human pose estimation from the 2D counterpart. On one hand, without explicit annotation, the 3D skeleton is hard to be accurately estimated from the occluded 2D pose. On the other hand, one occluded 2D pose might correspond to multiple 3D skeletons with low confidence parts. To address these issues, we decouple the 3D representation feature into view-invariant part termed occlusion-aware feature and view-dependent part termed rotation feature to facilitate subsequent optimization of the former. Then we propose an occlusion-aware contrastive representation based scheme (OCR-Pose) consisting of Topology Invariant Contrastive Learning module (TiCLR) and View Equivariant Contrastive Learning module (VeCLR). Specifically, TiCLR drives invariance to topology transformation, i.e., bridging the gap between an occluded 2D pose and the unoccluded one. While VeCLR encourages equivariance to view transformation, i.e., capturing the geometric similarity of the 3D skeleton in two views. Both modules optimize occlusion-aware constrastive representation with pose filling and lifting networks via an iterative training strategy in an end-to-end manner. OCR-Pose not only achieves superior performance against state-of-the-art unsupervised methods on unoccluded benchmarks, but also obtains significant improvements when occlusion is involved. Our project is available at https://sites.google.com/view/ocr-pose. Zhenbo Yu, Zhengyan Tong, Jinxian Liu, Wenjun Zhang 0001 |
ACM Multimedia | 6 |
| 2022 | Enhanced Preamble Based MAC Mechanism for IIoT-oriented PLC NetworkabstractIn this paper, we propose an enhanced preamble based media access control mechanism (E-PMAC), which can be applied in power line communication (PLC) network for Industrial Internet of Things (IIoT). We introduce detailed technologies used in E-PMAC, including delay calibration mechanism, preamble design, and slot allocation algorithm. With these technologies, E-PMAC is more robust than existing preamble based MAC mechanism (P-MAC). Besides, we analyze the disadvantage of P-MAC in multi-layer networking and design the networking process of E-PMAC to accelerate networking process. We analyze the complexity of networking process in P-MAC and E-PMAC and prove that E-PMAC has lower complexity than P-MAC. Finally, we simulate the single-layer networking and multi-layer networking of E-PMAC, P-MAC, and existing PLC protocol, i.e., IEEE1901.1. The simulation results indicate that E-PMAC spends much less time in networking than IEEE1901.1 and P-MAC. Finally, with our work, a PLC network based on E-PMAC mechanism can be realized. Biqian Feng, Yongpeng Wu 0001, Wenjun Zhang 0001 |
VTC Spring | 4 |
| 2022 | L0 structure-prior assisted blur-intensity aware efficient video deblurring
Chen Li 0021, Li Song 0001, Rong Xie 0004, Wenjun Zhang 0001 |
Neurocomputing | 4 |
| 2022 | Joint Device Detection, Channel Estimation, and Data Decoding With Collision Resolution for MIMO Massive Unsourced Random AccessabstractIn this paper, we investigate a joint device activity detection (DAD), channel estimation (CE), and data decoding (DD) algorithm for multiple-input multiple-output (MIMO) massive unsourced random access (URA). Different from the state-of-the-art slotted transmission scheme, the data in the proposed framework is split into only two parts. A portion of the data is coded by compressed sensing (CS) and the rest is low-density-parity-check (LDPC) coded. In addition to being part of the data, information bits in the CS phase also undertake the task of interleaving pattern design and CE. The principle of interleave-division multiple access (IDMA) is exploited to reduce the interference among devices in the LDPC phase. Based on the belief propagation (BP) algorithm, a low-complexity iterative message passing (MP) algorithm is utilized to decode the data embedded in these two phases separately. Moreover, combined with successive interference cancellation (SIC), the proposed joint DAD-CE-DD algorithm is performed to further improve performance by utilizing the belief of each other. Additionally, based on the energy detection (ED) and sliding window protocol (SWP), we develop a collision resolution protocol to handle the codeword collision, a common issue in the URA system. In addition to the complexity reduction, the proposed algorithm exhibits a substantial performance enhancement compared to the state-of-the-art in terms of efficiency and accuracy. Tianya Li, Yongpeng Wu 0001, Mengfan Zheng, Wenjun Zhang 0001, Chengwen Xing, Jianping An, Xiang-Gen Xia 0001, Chengshan Xiao |
IEEE J. Sel. Areas Commun. | 4 |
| 2022 | Random Access With Massive MIMO-OTFS in LEO Satellite CommunicationsabstractThis paper considers the joint channel estimation and device activity detection in the grant-free random access systems, where a large number of Internet-of-Things devices intend to communicate with a low-earth orbit satellite in a sporadic way. In addition, the massive multiple-input multiple-output (MIMO) with orthogonal time-frequency space (OTFS) modulation is adopted to combat the dynamics of the terrestrial-satellite link. We first analyze the input-output relationship of the single-input single-output OTFS when the large delay and Doppler shift both exist, and then extend it to the grant-free random access with massive MIMO-OTFS. Next, by exploring the sparsity of channel in the delay-Doppler-angle domain, a two-dimensional pattern coupled hierarchical prior with the sparse Bayesian learning and covariance-free method (TDSBL-FM) is developed for the channel estimation. Then, the active devices are detected by computing the energy of the estimated channel. Finally, the generalized approximate message passing algorithm combined with the sparse Bayesian learning and two-dimensional convolution (ConvSBL-GAMP) is proposed to decrease the computations of the TDSBL-FM algorithm. Simulation results demonstrate that the proposed algorithms outperform conventional methods. Boxiao Shen, Yongpeng Wu 0001, Jianping An, Chengwen Xing, Lian Zhao, Wenjun Zhang 0001 |
IEEE J. Sel. Areas Commun. | 6 |
| 2022 | Fine-Grained Video Captioning via Graph-based Multi-Granularity Interaction LearningabstractLearning to generate continuous linguistic descriptions for multi-subject interactive videos in great details has particular applications in team sports auto-narrative. In contrast to traditional video caption, this task is more challenging as it requires simultaneous modeling of fine-grained individual actions, uncovering of spatio-temporal dependency structures of frequent group interactions, and then accurate mapping of these complex interaction details into long and detailed commentary. To explicitly address these challenges, we propose a novel framework Graph-based Learning for Multi-Granularity Interaction Representation (GLMGIR) for fine-grained team sports auto-narrative task. A multi-granular interaction modeling module is proposed to extract among-subjects' interactive actions in a progressive way for encoding both intra- and inter-team interactions. Based on the above multi-granular representations, a multi-granular attention module is developed to consider action/event descriptions of multiple spatio-temporal resolutions. Both modules are integrated seamlessly and work in a collaborative way to generate the final narrative. In the meantime, to facilitate reproducible research, we collect a new video dataset from YouTube.com called Sports Video Narrative dataset (SVN). It is a novel direction as it contains 6K team sports videos (i.e., NBA basketball games) with 10K ground-truth narratives(e.g., sentences). Furthermore, as previous metrics such as METEOR (i.e., used in coarse-grained video caption task) DO NOT cope with fine-grained sports narrative task well, we hence develop a novel evaluation metric named Fine-grained Captioning Evaluation (FCE), which measures how accurate the generated linguistic description reflects fine-grained action details as well as the overall spatio-temporal interactional structure. Extensive experiments on our SVN dataset have demonstrated the effectiveness of the proposed framework for fine-grained team sports video auto-narrative. Yichao Yan, Ning Zhuang, Bingbing Ni, Jian Zhang 0079, Qi Tian 0001, Yi Xu 0001, Xiaokang Yang 0001, Wenjun Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 12 |
| 2022 | Mobile Communications, Computing, and Caching Resources Allocation for Diverse Services via Multi-Objetive Proximal Policy OptimizationabstractMobile services are becoming more diverse, making them have different demands on communications, computing, and caching (3C) resources in mobile systems. Unlike the traditional work that considers only one type of service, this paper designs a unified framework to characterize the different kinds of services, and jointly optimizes the 3C resources of the base station (BS) and mobile devices to provide differentiated quality of service (QoS) for diverse services. In the proposed framework, we model the task required by the mobile device to be generated at the BS, the mobile device, or both of them, which means the requested tasks are served through different paths, consuming different bandwidth, computing and caching resources. Since diverse services have different QoS, we formulate a multi-objective programming (MOP) to optimize the allocation of the 3C resources for minimizing the total delay while maximizing the number of executed tasks requested by the mobile devices. We transform the MOP problem as a multi-objective Markov decision process (MO-MDP) and design a multi-objective proximal policy optimization (MO-PPO) algorithm to solve the MO-MDP. The proposed MO-PPO first trains two sub-policies separately for the two objectives, and then combines them to search for Pareto dominating solutions. By alternately perform the separate training and the combination, we can finally obtain a set of Pareto optimal solutions and the corresponding Pareto front. Simulation results show that the proposed MO-PPO outperforms traditional methods in finding a higher-quality set of Pareto optimal solutions and can more appropriately allocate 3C resources to different types of services. Zhiyong Chen 0002, Benshun Yin, Yingjiao Li, Meixia Tao, Wenjun Zhang 0001 |
IEEE Trans. Commun. | 6 |
| 2022 | Massive Unsourced Random Access: Exploiting Angular Domain SparsityabstractThis paper investigates the unsourced random access (URA) scheme to accommodate numerous machine-type users communicating to a base station equipped with multiple antennas. Existing works adopt a slotted transmission strategy to reduce system complexity; they operate under the framework of coupled compressed sensing (CCS) which concatenates an outer tree code to an inner compressed sensing code for slot-wise message stitching. We suggest that by exploiting the MIMO channel information in the angular domain, redundancies required by the tree encoder/decoder in CCS can be removed to improve spectral efficiency, thereby an uncoupled transmission protocol is devised. To perform activity detection and channel estimation, we propose an expectation-maximization-aided generalized approximate message passing algorithm with a Markov random field support structure, which captures the inherent clustered sparsity structure of the angular domain channel. Then, message reconstruction in the form of a clustering decoder is performed by recognizing slot-distributed channels of each active user based on similarity. We put forward the slot-balanced$ K $-means algorithm as the kernel of the clustering decoder, resolving constraints and collisions specific to the application scene. Extensive simulations reveal that the proposed scheme achieves a better error performance at high spectral efficiency compared to the CCS-based URA schemes. Xinyu Xie, Yongpeng Wu 0001, Jianping An, Junyuan Gao, Wenjun Zhang 0001, Chengwen Xing, Kai-Kit Wong, Chengshan Xiao |
IEEE Trans. Commun. | 5 |
| 2022 | Thinking Inside Uncertainty: Interest Moment Perception for Diverse Temporal GroundingabstractGiven a language query, temporal grounding task is to localize temporal boundaries of the described event in an untrimmed video. There is a long-standing challenge that multiple moments may be associated with one same video-query pair, termed label uncertainty. However, existing methods struggle to localize diverse moments due to the lack of multi-label annotations. In this paper, we propose a novel Diverse Temporal Grounding framework (DTG) to achieve diverse moment localization with only single-label annotations. By delving into the label uncertainty, we find the diverse moments retrieved tend to involve similar actions/objects, driving us to perceive these interest moments. Specifically, we construct soft multi-label through semantic similarity of multiple video-query pairs. These soft labels reveal whether multiple moments in the intra-videos contain similar verbs/nouns, thereby guiding interest moment generation. Meanwhile, we put forward a diverse moment regression network (DMRNet) to achieve multiple predictions in a single pass, where plausible moments are dynamically picked out from the interest moments for joint optimization. Moreover, we introduce new metrics that better reveal multi-output performance. Extensive experiments conducted on Charades-STA and ActivityNet Captions show that our method achieves state-of-the-art performance in terms of both standard and new metrics. Hao Zhou 0014, Yan Luo 0003, Chuanping Hu, Wenjun Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | HazDesNet: An End-to-End Network for Haze Density PredictionabstractVision-based intelligent systems such as driver assistance systems and transportation systems should take into account weather conditions. The presence of haze in images can be a critical threat to driving scenarios. Haze density measures the visibility and usability of hazy images captured in real-world conditions. The prediction of haze density can be valuable in various vision-based intelligent systems, especially in those systems deployed in outdoor environments. Haze density prediction is a challenging task since the haze and many scene contents have a lot in common in appearance. Existing methods generally utilize different priors and design complex handcrafted features to predict the visibility or haze density of the image. In this article, we propose a novel end-to-end convolutional neural network (CNN) based method to predict haze density, named as HazDesNet. Our HazDesNet takes a hazy image as input and predicts a pixel-level haze density map. The density map is then refined and smoothed, and the average of the refined map is calculated as the global haze density of the image. To verify the performance of HazDesNet, a subjective human study is performed to build a Human Perceptual Haze Density (HPHD) database, which includes 500 real-world hazy images and 100 synthetic hazy images, and the corresponding human-rated perceptual haze density scores. Experimental results show that our method achieves the best haze density prediction performance on our built HPHD database and existing databases. Besides the global quantitative results, our HazDesNet is capable of predicting a continuous, stable, fine, and high-resolution haze density map. We will make the database and code publicly available athttps://github.com/JiaheZhang/HazDesNet. Xiongkuo Min, Yucheng Zhu, Guangtao Zhai, Jiantao Zhou 0001, Xiaokang Yang 0001, Wenjun Zhang 0001 |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2022 | QoE Driven VR 360° Video Massive MIMO TransmissionabstractMassive multiple-input and multiple-output (MIMO) enables ultra-high throughput and low latency for tile-based adaptive virtual reality (VR) 360° video transmission in wireless network. In this paper, we consider a massive MIMO system where multiple users in a single-cell theater watch an identical VR 360° video. Based on tile prediction, base station (BS) deliveries the tiles in predicted field of view (FoV) to users. By introducing practical supplementary transmission for missing tiles and unacceptable VR sickness, we propose the first stable transmission scheme for VR video. we formulate an integer non-linear programming (INLP) problem to maximize users’ average quality of experience (QoE) score. Moreover, we derive the achievable spectral efficiency (SE) expression of predictive tile groups and the approximately achievable SE expression of missing tile groups, respectively. Analytically, the overall throughput is related to the number of tile groups and the length of pilot sequences. By exploiting the relationship between the structure of viewport tiles and SE expression, we propose a multi-lattice multi-stream grouping method aimed at improving the overall throughput for VR video transmission. Moreover, we analyze the relationship between QoE objective and number of predictive tile. We transform the original INLP problem into an integer linear programming problem by setting the predictive tiles groups as some constants. With variable relaxation and recovery, we obtain the optimal average QoE. Extensive simulation results validate that the proposed algorithm effectively improves QoE. Guangtao Zhai, Yongpeng Wu 0001, Xiongkuo Min, Wenjun Zhang 0001, Zhi Ding 0001, Chengshan Xiao |
IEEE Trans. Wirel. Commun. | 5 |
| 2021 | Progressive Stage-Wise Learning for Unsupervised Feature Representation EnhancementabstractUnsupervised learning methods have recently shown their competitiveness against supervised training. Typically, these methods use a single objective to train the en-tire network. But one distinct advantage of unsupervised over supervised learning is that the former possesses more variety and freedom in designing the objective. In this work, we explore new dimensions of unsupervised learning by proposing the Progressive Stage-wise Learning (PSL) framework. For a given unsupervised task, we design multi-level tasks and define different learning stages for the deep network. Early learning stages are forced to focus on low-level tasks while late stages are guided to extract deeper information through harder tasks. We discover that by progressive stage-wise learning, unsupervised feature representation can be effectively enhanced. Our extensive experiments show that PSL consistently improves results for the leading unsupervised learning methods. Zefan Li, Chenxi Liu 0001, Alan L. Yuille, Bingbing Ni, Wenjun Zhang 0001, Wen Gao 0001 |
CVPR | 5 |
| 2021 | 3D Human Action Representation Learning via Cross-View Consistency PursuitabstractIn this work, we propose a Cross-view Contrastive Learning framework for unsupervised 3D skeleton-based action Representation (CrosSCLR), by leveraging multi-view complementary supervision signal. CrosSCLR consists of both single-view contrastive learning (Skeleton-CLR) and cross-view consistent knowledge mining (CVC-KM) modules, integrated in a collaborative learning manner. It is noted that CVC-KM works in such a way that high-confidence positive/negative samples and their distributions are exchanged among views according to their embedding similarity, ensuring cross-view consistency in terms of contrastive context, i.e., similar distributions. Extensive experiments show that CrosSCLR achieves remarkable action recognition results on NTU-60 and NTU-120 datasets under unsupervised settings, with observed higher-quality action representations. Our code is available at https://github.com/LinguoLi/CrosSCLR. Linguo Li, Minsi Wang, Bingbing Ni, Jiancheng Yang, Wenjun Zhang 0001 |
CVPR | 6 |
| 2021 | Detection of Abrupt Change in Channel Covariance Matrix for Multi-Antenna CommunicationabstractThe knowledge of channel covariance matrices is of paramount importance to the estimation of instantaneous channels and the design of beamforming vectors in multi-antenna systems. In practice, an abrupt change in channel covariance matrices may occur due to the change in the environment and the user location. Although several works have proposed efficient algorithms to estimate the channel covariance matrices after any change occurs, how to detect such a change accurately and quickly is still an open problem in the literature. In this paper, we focus on channel covariance change detection between a multi-antenna base station (BS) and a single-antenna user equipment (UE). To provide theoretical performance limit, we first propose a genie-aided change detector based on the log-likelihood ratio (LLR) test assuming the channel covariance matrix after change is known, and characterize the corresponding missed detection and false alarm probabilities. Then, this paper considers the practical case where the channel covariance matrix after change is unknown. The maximum likelihood (ML) estimation technique is used to predict the covariance matrix based on the received pilot signals over a certain number of coherence blocks, building upon which the LLR-based change detector is employed. Numerical results show that our proposed scheme can detect the change with low error probability even when the number of channel samples is small such that the estimation of the covariance matrix is not that accurate. This result verifies the possibility to detect the channel covariance change both accurately and quickly in practice. Runnan Liu, Liang Liu 0003, Dazhi He, Wenjun Zhang 0001, Erik G. Larsson |
GLOBECOM | 4 |
| 2021 | Geometric Granularity Aware Pixel-to-MeshabstractPixel-to-mesh has wide applications, especially in virtual or augmented reality, animation and game industry. However, existing mesh reconstruction models perform unsatisfactorily in local geometry details due to ignoring mesh topology information during learning. Besides, most methods are constrained by the initial template, which cannot reconstruct meshes of various genus. In this work, we propose a geometric granularity-aware pixel-to-mesh framework with a fidelity-selection-and-guarantee strategy, which explicitly addresses both challenges. First, a geometry structure extractor is proposed for detecting local high structured parts and capturing local spatial feature. Second, we apply it to facilitate pixel-to-mesh mapping and resolve coarse details problem caused by the neglect of structural information in previous practices. Finally, a mesh edit module is proposed to encourage non-zero genus topology to emergence by fine-grained topology modification and a patching algorithm is introduced to repair the non-closed boundaries. Extensive experimental results, both quantitatively and visually have demonstrated the high reconstruction fidelity achieved by the proposed framework. Bingbing Ni, Jinxian Liu, Dingyi Rong, Ye Qian, Wenjun Zhang 0001 |
ICCV | 6 |
| 2021 | Towards Alleviating the Modeling Ambiguity of Unsupervised Monocular 3D Human Pose EstimationabstractIn this work, we study the ambiguity problem in the task of unsupervised 3D human pose estimation from 2D counterpart. On one hand, without explicit annotation, the scale of 3D pose is difficult to be accurately captured (scale ambiguity). On the other hand, one 2D pose might correspond to multiple 3D gestures, where the lifting procedure is inherently ambiguous (pose ambiguity). Previous methods generally use temporal constraints (e.g., constant bone length and motion smoothness) to alleviate the above issues. However, these methods commonly enforce the outputs to fulfill multiple training objectives simultaneously, which often lead to sub-optimal results. In contrast to the majority of previous works, we propose to split the whole problem into two sub-tasks, i.e., optimizing 2D input poses via a scale estimation module and then mapping optimized 2D pose to 3D counterpart via a pose lifting module. Furthermore, two temporal constraints are proposed to alleviate the scale and pose ambiguity respectively. These two modules are optimized via a iterative training scheme with corresponding temporal constraints, which effectively reduce the learning difficulty and lead to better performance. Results on the Human3.6M dataset demonstrate that our approach improves upon the prior art by 23.1% and also outperforms several weakly supervised approaches that rely on 3D annotations. Our project is available at https://sites.google.com/view/ambiguity-aware-hpe. Zhenbo Yu, Bingbing Ni, Jingwei Xu 0005, Chenglong Zhao, Wenjun Zhang 0001 |
ICCV | 6 |
| 2021 | Skeleton2Mesh: Kinematics Prior Injected Unsupervised Human Mesh RecoveryabstractIn this paper, we decouple unsupervised human mesh recovery into the well-studied problems of unsupervised 3D pose estimation, and human mesh recovery from estimated 3D skeletons, focusing on the latter task. The challenges of the latter task are two folds: (1) pose failure (i.e., pose mismatching – different skeleton definitions in dataset and SMPL , and pose ambiguity – endpoints have arbitrary joint angle configurations for the same 3D joint coordinates). (2) shape ambiguity (i.e., the lack of shape constraints on body configuration). To address these issues, we propose Skeleton2Mesh, a novel lightweight framework that recovers human mesh from a single image. Our Skeleton2Mesh contains three modules, i.e., Differentiable Inverse Kinematics (DIK), Pose Refinement (PR) and Shape Refinement (SR) modules. DIK is designed to transfer 3D rotation from estimated 3D skeletons, which relies on a minimal set of kinematics prior knowledge. Then PR and SR modules are utilized to tackle the pose ambiguity and shape ambiguity respectively. All three modules can be incorporated into Skeleton2Mesh seamlessly via an end-to-end manner. Furthermore, we utilize an adaptive joint regressor to alleviate the effects of skeletal topology from different datasets. Results on the Human3.6M dataset for human mesh recovery demonstrate that our method improves upon the previous unsupervised methods by 32.6% under the same setting. Qualitative results on in-the-wild datasets exhibit that the recovered 3D meshes are natural, realistic. Our project is available at https://sites.google.com/view/skeleton2mesh. Zhenbo Yu, Jingwei Xu 0005, Bingbing Ni, Chenglong Zhao, Minsi Wang, Wenjun Zhang 0001 |
ICCV | 7 |
| 2021 | SVM Based Fast CU Partitioning Algorithm for VVC Intra CodingabstractRecently, Joint Video Experts Team (JVET) has completed the new Versatile Video Coding (H.266/VVC) standard. VVC employs a new block partition structure named quad-tree with nested multi-type tree (QTMT) to improve coding efficiency. However, the new block partition structure increases huge encoding time compared with HEVC for brute-force ratedistortion (RD) optimization. To reduce encoding complexity, we propose a Support Vector Machine (SVM) based fast CU partitioning algorithm for VVC intra coding in this paper which terminates redundant partitions early by predicting the partition of CU using texture information. We trained classifiers for CUs of different sizes to improve accuracy and control the complexity of the classifiers themselves. Different thresholds are set for each classifier to achieve a trade-off between encoding complexity and RD performance. Experimental results show that the proposed method can save encoder time ranging from 30.78% to 63.16% with 1.10% to 2.71% BD-BR increase. Yan Huang 0033, Li Song 0001, Wenjun Zhang 0001 |
ISCAS | 5 |
| 2021 | Optimizing Channel Estimation Overhead for OTFS with Prior Channel StatisticsabstractThe recently proposed orthogonal time-frequency space (OTFS) modulation scheme is able to provide significant performance gain over orthogonal frequency division multiplexing (OFDM) in high Doppler spread scenarios. Qualified channel estimation in delay-Doppler domain is the prerequisite for such good performance, but it requires a large number of guard and pilot symbols, which degrades channel capacity significantly. In this paper, a prior channel statistics based scheme is proposed to maximize the system ergodic capacity by optimizing the channel estimation overhead while ensuring the high-quality performance of the OTFS over delay-Doppler channels. We first investigate the signal-to-interference-plus-noise ratio (SINR) performance of the proposed scheme and derive the closed-form ergodic capacity of the OTFS system on its basis. And then, the capacity maximization problem is studied given the root-mean-square (RMS) delay spread rather than a request for instantaneous perfect channel state information (CSI). In addition, we reveal that the additive white Gaussian noise (AWGN) power has an impact on the solution to the optimization problem, so optimization instances are proposed to treat differently according to SNR levels. Extensive simulations are carried out to evaluate the performance of the proposed overhead reduction scheme. Numerical results demonstrate the superiority of the proposed scheme, and the effect of channel characteristics and system parameters are also revealed. Runnan Liu, Dazhi He, Yin Xu 0001, Wenjun Zhang 0001 |
WCNC | 5 |
| 2021 | Uplink transmission design for crowded correlated cell-free massive MIMO-OFDM systems
Junyuan Gao, Yongpeng Wu 0001, Wenjun Zhang 0001, Fan Wei 0004 |
Sci. China Inf. Sci. | 4 |
| 2021 | Modeling Acceleration Properties for Flexible INTRA HEVC Complexity ControlabstractIt is a very well-known fact, that the high complexity of the High Efficiency Video Coding standard (HEVC) is the main hurdle for its wide deployment and use. To tackle this problem, a number of recent research outcomes exploit heuristic algorithms and machine learning, including deep learning, to reduce the coding complexity. However, in most cases, each encoder module, i.e., encoding process, is first accelerated individually, and then different acceleration algorithms are manually combined. Without a holistic strategy, the acceleration potential of multi-module combination is not exploited and the Rate-Distortion (RD) loss is generally not well controlled. To tackle these shortcomings, this paper exploits the acceleration properties of different modules, i.e., the numerical representation of potential time saving and possible RD loss, from which a heuristic model is explored. Then a Heuristic Model Oriented Framework (HMOF) is proposed which adapts the properties of modules to underlying acceleration algorithms. In the framework, two advanced acceleration algorithms, including Border Considered CNN (BC-CNN)-based Coding Unit (CU) partition and Naive Bayes-based Prediction Unit (PU) partition, are proposed for the CU and PU modules, respectively. Further, by leveraging the heuristic model as the guidance to combine the proposed acceleration algorithms, HMOF is globally optimized, where different time saving budgets are wisely allocated to different modules and a theoretically minimal RD loss is achieved. According to the experimental results, through fusing a suitable deep learning technique and a Bayes-Based prediction, the proposed acceleration framework HMOF enable multiple acceleration choices. Here the proposed joint optimization strategy help to make a choice leading to the best cost-performance. Furthermore, within the proposed framework, intra coding time can be precisely controlled with negligible Bjøntegaard delta bit-rate (BDBR) loss. In this context, as a complexity control method, HMOF outperforms the state-of-the-art complexity reduction algorithms under a similar complexity reduction ratio. These results partially demonstrate the superiority of the proposed technique. Yan Huang 0033, Li Song 0001, Rong Xie 0004, Ebroul Izquierdo, Wenjun Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2021 | Learned Resolution Scaling Powered Gaming-as-a-Service at ScaleabstractBuilt on the explosive advancement of cloud and telecommunication technologies, Gaming-as-a-Service (GaaS) or cloud gaming system is expected to revolutionize the traditional multi-billion video game market in the near future. This wave is analogous to the rise of live-video-streaming-based-Netflix to replace conventional DVD rental business for movies and TVs. In practice, a successful GaaS platform need to operate in a transparent mode without requiring substantial efforts from both content providers and end users, and offer the pristine quality of experience (QoE) at an affordable cost. Our analysis suggests that GaaS provisioning cost can be reduced significantly by enforcing the game video rendering and streaming at a lower resolution (so as to increase the user concurrency in the cloud and reduce the streaming bandwidth over the network). However, streaming video at a lower resolution may deteriorate the QoE. To maintain the client QoE at the level using the default-native resolution for streaming or even enhance it, we introduce the learned resolution scaling (LRS), which leverages the computational capabilities at clients/edges to restore/improve the reconstructed image/video quality via stacked deep neural networks (DNN). We integrate this LRS into a commercialized GaaS platform - AnyGame, to study its efficiency and complexity quantitatively. Extensive real-life experiments have shown that LRS-powered AnyGame offers the state-of-the-art performance, and the lower operational cost, paving the road for a potential success of GaaS over the Internet. Additionally, we dive into proposed LRS via ablation studies to further demonstrate its consistent performance, including the discussions on trade-off between efficiency and complexity, alternative training sets, etc. Hao Chen 0036, Ming Lu 0003, Zhan Ma 0001, Xu Zhang 0006, Yiling Xu, Qiu Shen, Wenjun Zhang 0001 |
IEEE Trans. Multim. | 7 |
| 2021 | Group Re-Identification With Group Context Graph Neural NetworksabstractGroup re-identification aims to match groups of people across disjoint cameras. In this task, the contextual information from neighbor individuals can be exploited for re-identifying each individual within the group as well as the entire group. However, compared with single person re-identification, it brings new challenges including group layout and group membership changes. Motivated by the observation that individuals who are close together are more likely to keep in the same group under different cameras than those who are far apart, we propose to model each group as a spatial K-nearest neighbor graph (SKNNG) and design a group context graph neural network (GCGNN) for graph representation learning. Specifically, for each node in the graph, the proposed GCGNN learns an embedding which aggregates the contextual information from neighbor nodes. We design multiple weighting kernels for neighborhood aggregation based on the graph properties including node in-degrees and spatial relationship attributes. We compute the similarity scores between node embeddings of two graphs for group member association and obtain the matching score between the two graphs by summing up the similarity scores of all linked node pairs. Experimental results on three public datasets show that our approach performs favorably against state-of-the-art methods and achieves high efficiency. Ji Zhu 0002, Hua Yang 0001, Weiyao Lin, Nian Liu 0002, Jia Wang 0004, Wenjun Zhang 0001 |
IEEE Trans. Multim. | 6 |
| 2020 | Adversarial Domain Adaptation with Domain MixupabstractRecent works on domain adaptation reveal the effectiveness of adversarial learning on filling the discrepancy between source and target domains. However, two common limitations exist in current adversarial-learning-based methods. First, samples from two domains alone are not sufficient to ensure domain-invariance at most part of latent space. Second, the domain discriminator involved in these methods can only judge real or fake with the guidance of hard label, while it is more reasonable to use soft scores to evaluate the generated images or features, i.e., to fully utilize the inter-domain information. In this paper, we present adversarial domain adaptation with domain mixup (DM-ADA), which guarantees domain-invariance in a more continuous latent space and guides the domain discriminator in judging samples' difference relative to source and target domains. Domain mixup is jointly conducted on pixel and feature level to improve the robustness of models. Extensive experiments prove that the proposed approach can achieve superior performance on tasks with various degrees of domain shift and data complexity. Jian Zhang 0079, Bingbing Ni, Teng Li 0001, Chengjie Wang 0001, Qi Tian 0001, Wenjun Zhang 0001 |
AAAI | 7 |
| 2020 | FACT: Fused Attention for Clothing Transfer with Generative Adversarial NetworksabstractClothing transfer is a challenging task in computer vision where the goal is to transfer the human clothing style in an input image conditioned on a given language description. However, existing approaches have limited ability in delicate colorization and texture synthesis with a conventional fully convolutional generator. To tackle this problem, we propose a novel semantic-based Fused Attention model for Clothing Transfer (FACT), which allows fine-grained synthesis, high global consistency and plausible hallucination in images. Towards this end, we incorporate two attention modules based on spatial levels: (i) soft attention that searches for the most related positions in sentences, and (ii) self-attention modeling long-range dependencies on feature maps. Furthermore, we also develop a stylized channel-wise attention module to capture correlations on feature levels. We effectively fuse these attention modules in the generator and achieve better performances than the state-of-the-art method on the DeepFashion dataset. Qualitative and quantitative comparisons against the baselines demonstrate the effectiveness of our approach. Li Song 0001, Rong Xie 0004, Wenjun Zhang 0001 |
AAAI | 5 |
| 2020 | Cross-Domain Detection via Graph-Induced Prototype AlignmentabstractApplying the knowledge of an object detector trained on a specific domain directly onto a new domain is risky, as the gap between two domains can severely degrade model's performance. Furthermore, since different instances commonly embody distinct modal information in object detection scenario, the feature alignment of source and target domain is hard to be realized. To mitigate these problems, we propose a Graph-induced Prototype Alignment (GPA) framework to seek for category-level domain alignment via elaborate prototype representations. In the nutshell, more precise instance-level features are obtained through graph-based information propagation among region proposals, and, on such basis, the prototype representation of each class is derived for category-level domain alignment. In addition, in order to alleviate the negative effect of class-imbalance on domain adaptation, we design a Class-reweighted Contrastive Loss to harmonize the adaptation training process. Combining with Faster R-CNN, the proposed framework conducts feature alignment in a two-stage manner. Comprehensive results on various cross-domain detection tasks demonstrate that our approach outperforms existing methods with a remarkable margin. Our code is available at https://github.com/ChrisAllenMing/GPA-detection. Bingbing Ni, Qi Tian 0001, Wenjun Zhang 0001 |
CVPR | 5 |
| 2020 | Deep Kinematics Analysis for Monocular 3D Human Pose EstimationabstractFor monocular 3D pose estimation conditioned on 2D detection, noisy/unreliable input is a key obstacle in this task. Simple structure constraints attempting to tackle this problem, e.g., symmetry loss and joint angle limit, could only provide marginal improvements and are commonly treated as auxiliary losses in previous researches. Thus it still remains challenging about how to effectively utilize the power of human prior knowledge for this task. In this paper, we propose to address above issue in a systematic view. Firstly, we show that optimizing the kinematics structure of noisy 2D inputs is critical to obtain accurate 3D estimations. Secondly, based on corrected 2D joints, we further explicitly decompose articulated motion with human topology, which leads to more compact 3D static structure easier for estimation. Finally, temporal refinement emphasizing the validity of 3D dynamic structure is naturally developed to pursue more accurate result. Above three steps are seamlessly integrated into deep neural models, which form a deep kinematics analysis pipeline concurrently considering the static/dynamic structure of 2D inputs and 3D outputs. Extensive experiments show that proposed framework achieves state-of-the-art performance on two widely used 3D human action datasets. Meanwhile, targeted ablation study shows that each former step is critical for the latter one to obtain promising results. Jingwei Xu 0005, Zhenbo Yu, Bingbing Ni, Jiancheng Yang, Xiaokang Yang 0001, Wenjun Zhang 0001 |
CVPR | 6 |
| 2020 | Learning to Combine: Knowledge Aggregation for Multi-source Domain Adaptation
Bingbing Ni, Wenjun Zhang 0001 |
ECCV (8) | 4 |
| 2020 | Energy-efficiency of Massive Random Access with Individual CodebookabstractThe massive machine-type communication has been one of the most representative services for future wireless networks. It aims to support massive connectivity of user equipments (UEs) which sporadically transmit packets with small size. In this work, we assume the number of UEs grows linearly and unboundedly with blocklength and each UE has an individual codebook. Among all UEs, an unknown subset of UEs are active and transmit a fixed number of data bits to a base station over a shared -spectrum radio link. Under these settings, we derive the achievability and converse bounds on the minimum energy-per-bit for reliable random access over quasi-static fading channels with and without channel state information (CSI) at the receiver. These bounds provide energy-efficiency guidance for new schemes suited for massive random access. Simulation results indicate that the orthogonalization scheme TDMA is energy-inefficient for large values of UE density μ. Besides, the multi-user interference can be perfectly cancelled when μ is below a critical threshold. In the case of no-CSI, the energy-per-bit for random access is only a bit more than that with the knowledge UE activity. Junyuan Gao, Yongpeng Wu 0001, Wenjun Zhang 0001 |
GLOBECOM | 3 |
| 2020 | Massive Unsourced Random Access for Massive MIMO Correlated ChannelsabstractThis paper investigates the massive random access for a huge amount of user devices served by a base station (BS) equipped with a massive number of antennas. We consider a grant-free unsourced random access (U-RA) scheme where all users possess the same codebook and the BS aims at declaring a list of transmitted codewords and recovering the messages sent by active users. Most of the existing works concentrate on applying U-RA in the oversimplified independent and identically distributed (i.i.d.) channels. In this paper, we consider a fairly general joint-correlated MIMO channel model with line-of-sight components for the realistic outdoor wireless propagation environments. We conduct the activity detection for the emitted codewords by performing an improved coordinate descent approach with Bayesian learning automaton to solve a covariance-based maximum likelihood estimation problem. The proposed algorithm exhibits a faster convergence rate than traditional descent approaches. We further employ a coupled coding scheme to resolve the issue that the dimensions of the common codebook expand exponentially with user payload size in the practical massive machine-type communications scenario. Our simulations reveal that to achieve an error probability of 0.05 for reliable communications in correlated channels, one must pay a 0.9 to 1.3 dB penalty comparing to the minimum signal to noise ratio needed in i.i.d. channels on condition that a sufficient number of receiving antennas is equipped at the BS. Xinyu Xie, Yongpeng Wu 0001, Junyuan Gao, Wenjun Zhang 0001 |
GLOBECOM | 4 |
| 2020 | Realistic Talking Face Synthesis With Geometry-Aware Feature TransformationabstractRecent studies have shown remarkable success in synthesizing realistic talking faces by exploiting generative adversarial networks. However, existing methods are mostly target specific that cannot generate images of previously unseen people, and they suffer from artifacts such as blurriness and mismatching of facial details. In this paper, we tackle these problems by proposing a target-agnostic framework. We introduce a geometry-aware feature transformation module to achieve shape transfer while preserving the appearance of the source face. To further improve image quality of synthesized results, we present a multi-scale spatially-consistent transfer unit to maintain spatial consistency between the encoder and decoder features. Experimental results show that our model is able to synthesize photo-realistic talking faces which are previously unseen, outperforming state-of-the-art methods both qualitatively and quantitatively. Jun Ling, Li Song 0001, Rong Xie 0004, Wenjun Zhang 0001 |
ICIP | 5 |
| 2020 | Dynamic Spectrum Allocation by 5G Base StationabstractIn 5G era, the base stations are capable of providing multiple services in various scenarios (e.g. vehicle network, Internet of Things), which provides fine opportunity to enhance spectrum efficiency. Base stations can flexibly utilize the idle frequency band for spatiotemporal low-demand services and guarantee services with high priority (e.g. urgent broadcasting), which construct a distributed architecture for spectrum allocation. In this paper, we provide a dynamic spectrum allocation scheme in base station, which can flexibly rearrange spectrum considering service priority, energy consumption and renting cost. We use Lyapunov optimization method to solve the problem. Moreover, we propose online Lyapunov optimization algorithm (OLOA) to figure out the optimal solution of penalty-and-drift function and show mathematical proofs on the performance of the algorithm. The simulation results show that the superiority and stability of our algorithm, which corroborates theoretical analysis. Yizhe Zhang 0003, Dazhi He, Wen He 0001, Yin Xu 0001, Yunfeng Guan 0001, Wenjun Zhang 0001 |
IWCMC | 6 |
| 2020 | Resolution Booster: Global Structure Preserving Stitching Method for Ultra-High Resolution Image Translation
Siying Zhai, Xiwei Hu, Xuanhong Chen, Bingbing Ni, Wenjun Zhang 0001 |
MMM (1) | 5 |
| 2020 | Bandit Learning-based Service Placement and Resource Allocation for Mobile Edge ComputingabstractService placement is a significant issue in mobile edge computing (MEC) system. Many works have proposed efficient offline approaches for service placement problems in MEC system. However, because of the randomness and uncertainty of mobile networks, it is impractical for these approaches to be implemented. Facing these uncertainty, we propose an online service placement scheme for MEC system without knowing service demand and network states in advance. In order to maximize the long-term accumulated reward obtained by service placement with limited resource constraint, we analyse this problem by a combinatorial multi-armed bandit (MAB) framework. In addition, because we simultaneously consider the service placement and resource allocation among services, it can be formulated as a multiple choice knapsack problem (MCKP) in each time slot. To solve this long-term reward maximization problem, we first propose a combinatorial upper bound confidence(CUCB)-based online service placement and resource allocation scheme. Then, we analyse the performance of this algorithm theoretically. Finally, simulation results show the efficiency of the algorithm. Wen He 0001, Dazhi He, Yizhe Zhang 0003, Yin Xu 0001, Yunfeng Guan 0001, Wenjun Zhang 0001 |
PIMRC | 7 |
| 2020 | Polar Coding and Sparse Spreading for Massive Unsourced Random AccessabstractIn this paper, we propose a new polar coding scheme for the unsourced, uncoordinated Gaussian random access channel. Our scheme is based on sparse spreading, treat interference as noise and successive interference cancellation (SIC). On the transmitters side, each user randomly picks a code-length and a transmit power from multiple choices according to some probability distribution to encode its message, and an interleaver to spread its encoded codeword bits across the entire transmission block. The encoding configuration of each user is transmitted by compressive sensing, similar to some previous works. On the receiver side, after recovering the encoding configurations of all users, it applies single-user polar decoding and SIC to recover the message list. Numerical results show that our scheme outperforms all previous schemes for active user number Ka≥ 250, and provides competitive performance for Ka≤ 225. Moreover, our scheme has much lower complexity compared to other schemes as we only use single-user polar coding. Mengfan Zheng, Yongpeng Wu 0001, Wenjun Zhang 0001 |
VTC Fall | 3 |
| 2020 | Modeling the Perceptual Quality of Viewport Adaptive Omnidirectional Video StreamingabstractInstead of streaming the entire OmniDirectional Videos (ODVs) that are often sampled at ultra high definition and high frame rate, a viewport adaptive streaming is preferred in practice. We usually stream the High-Quality (HQ) content within current viewport, while Low-Quality (LQ) elsewhere to save the network bandwidth consumption. Such scheme would lead to a quality refinement after user adapts his/her focus to a new viewport. In this paper, we thus model the perceptual impact of the quality variations (through adjusting the Quantization Stepsize (QS or q) and Spatial Resolution (SR or s)) with respect to the Refinement Duration (RD or τ) when performing the refinement from an arbitrary LQ scale to an arbitrary HQ one. A number of quality variations are studied to cover sufficient use cases in practice, resulting in a unified analytical model, as a product of separable exponential functions that measure the QS and SR induced perceptual impacts in terms of the RD, and a perceptual index measuring the subjective quality of corresponding viewport video after refinement. This model is first validated in a managed lab environment via independent subjective assessments by constraining user's navigation to avoid unexpected noise, where both Pearson Correlation Coefficient (PCC) and Spearman's Rank Correlation Coefficient (SRCC) are around 0.97. We then extend the validations in a real-life viewport-dependent streaming system, still yielding PCC and SRCC about 0.96 when comparing collected subjective scores with model predictions. Shaowei Xie, Yiling Xu, Qiu Shen, Zhan Ma 0001, Wenjun Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2020 | Scale-Aware Crowd Counting via Depth-Embedded Convolutional Neural NetworksabstractScale variation of pedestrians in a crowd image presents a significant challenge for vision-based people counting systems. Such variations are mainly caused by perspective-related distortions due to the camera pose relative to the ground plane. Following the density-based counting paradigm, we postulate that generating density values adaptive to object scales plays a critical role in the accuracy of the final counting results. Motivated by this, we distill the underlying information from depth cues to obtain scale-aware representations that can respond to object scales considering the fact that the scale is inversely proportional to the object depth. Specifically, we propose a depth embedding module as add-ons into existing networks. This module exploits essential depth cues to spatially re-calibrate the magnitude of the original features. In this way, the objects, although in the same class, will attain distinct representations according to their scales, which directly benefits the estimation of scale-aware density values. We conduct a comprehensive analysis of the effects of the depth embedding module and validate that exploiting depth cues to perceive object scale variations in convolutional neural networks improves crowd counting performances. Our experiments demonstrate the effectiveness of the proposed approach on four popular benchmark datasets. Muming Zhao, Jian Zhang 0002, Fatih Porikli, Bingbing Ni, Wenjun Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2020 | MUGGLE: MUlti-Stream Group Gaze Learning and EstimationabstractBeing able to accurately predict the common gaze point of a group of persons is of particular interest to precise marketing and automatic group attention assessment. Group gaze estimation faces challenges including small face/head size and outlier observers. To address these challenges, we proposed a novel framework called Multi-stream Group Gaze Learning and Estimation (MUGGLE). The MUGGLE infrastructure includes two inference streams: 1) a holistic stream which utilizes fused attention map as input to a global deep convolutional structure to explore the global geometric configurations and contexts of interesting persons in the scene; and 2) an aggregative stream which robustly aggregates individual gazes via a recurrent structure (e.g., LSTM) to obtain outlier-tolerant estimation. Both streams are seamlessly integrated via a fusion network. Extensive experiments are performed on a fully annotated group gaze image dataset with 8,000+ images and 100,000+ faces (which is publicly releasable). The results demonstrate the effectiveness of the proposed MUGGLE framework in group gaze estimation. Ning Zhuang, Bingbing Ni, Yi Xu 0001, Xiaokang Yang 0001, Wenjun Zhang 0001, Zefan Li, Wen Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2020 | Efficient Mobile Video Streaming via Context-Aware RaptorQ-Based Unequal Error ProtectionabstractMobile video streaming systems typically apply the forward error correction (FEC) at the application layer to cope with packet-level transmission errors, which complements the bit-level correction mechanisms at the physical layer. However, most existing works fail to exploit the block-level dependencies in both intra and interframe coding modes of a single-layer compressed video, and thus are less efficient for the prevailing H.264/AVC and/or H.265/HEVC compatible single-layer video application. To this end, we propose a low-complexity FEC, i.e., context-aware RaptorQ (CA-RQ) with unequal error protection (UEP), to improve the error recovery performance of the singlelayer mobile video streaming, through incorporating the blocklevel dependencies in the compressed video data. We use a packet-level video transmission distortion model that considers the dependencies in both spatial and temporal domains, to quantify the importance of video packets within a group of pictures (GoP). The compressed video packets are categorized and grouped into several classes according to their importance to construct the CA-RQ code with the UEP property. We provide a theoretical analysis on redundancy allocation bounds to demonstrate the superior performance of proposed CA-RQ over the standard RaptorQ code. In the meantime, extensive simulations have shown that our scheme not only offers much better subjective visual quality with less than 50% additional redundant symbols as compared to the Macroblock-Based UEP (MB-UEP) scheme, but also outperforms the MB-UEP and classical equal error protection (EEP)-based schemes, by a 0.45%'5.71% and 0.94%'6.78% margin, respectively, in reconstructed quality evaluated using the structural similarity (SSIM) index, across a reasonable range of redundancy proportions. Hao Chen 0036, Xu Zhang 0006, Yiling Xu, Zhan Ma 0001, Wenjun Zhang 0001 |
IEEE Trans. Multim. | 5 |
| 2019 | Variational Convolutional Neural Network PruningabstractWe propose a variational Bayesian scheme for pruning convolutional neural networks in channel level. This idea is motivated by the fact that deterministic value based pruning methods are inherently improper and unstable. In a nutshell, variational technique is introduced to estimate distribution of a newly proposed parameter, called channel saliency, based on this, redundant channels can be removed from model via a simple criterion. The advantages are two-fold: 1) Our method conducts channel pruning without desire of re-training stage, thus improving the computation efficiency. 2) Our method is implemented as a stand-alone module, called variational pruning layer, which can be straightforwardly inserted into off-the-shelf deep learning packages, without any special network design. Extensive experimental results well demonstrate the effectiveness of our method: For CIFAR-10, we perform channel removal on different CNN models up to 74\% reduction, which results in significant size reduction and computation saving. For ImageNet, about 40% channels of ResNet-50 are removed without compromising accuracy. Chenglong Zhao, Bingbing Ni, Jian Zhang 0079, Qiwei Zhao, Wenjun Zhang 0001, Qi Tian 0001 |
CVPR | 5 |
| 2019 | Leveraging Heterogeneous Auxiliary Tasks to Assist Crowd CountingabstractCrowd counting is a challenging task in the presence of drastic scale variations, the clutter background, and severe occlusions, etc. Existing CNN-based counting methods tackle these challenges mainly by fusing either multi-scale or multi-context features to generate robust representations. In this paper, we propose to address these issues by leveraging the heterogeneous attributes compounded in the density map. We identify three geometric/semantic/numeric attributes essentially important to the density estimation, and demonstrate how to effectively utilize these heterogeneous attributes to assist the crowd counting by formulating them into multiple auxiliary tasks. With the multi-fold regularization effects induced by the auxiliary tasks, the backbone CNN model is driven to embed desired properties explicitly and thus gains robust representations towards more accurate density estimation. Extensive experiments on three challenging crowd counting datasets have demonstrated the effectiveness of the proposed approach. Muming Zhao, Jian Zhang 0002, Wenjun Zhang 0001 |
CVPR | 4 |
| 2019 | Latency Minimization for Full-Duplex Mobile-Edge Computing SystemabstractMobile-edge computing (MEC) which employs cloud computing platforms at the network edge is an emerging paradigm for 5G networks. Users can be provided with lower latency and energy consumption by offloading computation tasks to the edge cloud. However, it is hard for MEC to guarantee low latency when large amount of users share the limited spectrum resource to offload computation tasks or download computation results because of high transmission delay. Full-duplex (FD) communication that allows simultaneous transmission and reception of signals over the same frequency band, is a promising solution to the shortage of spectrum resource. In this paper, we investigate a novel multi-user FD-MEC system involving both offloading and downloading processes. With the help of the FD capable BS, two half-duplex (HD) users can form a FD pair which can share the same time slots and frequency band for uplink and downlink transmission. To minimize the completion time of all users in the system, we formulate a joint optimization problem of time, power and user pairing scheme, which is a mixed-integer nonlinear programing (MINLP). This problem is further divided into two layers. For the inner layer problem, we obtain the optimal solution by a bisection method. While for the outer layer problem, we use concave-convex procedure (CCCP) to transform it into a tractable form and obtain a stationary point. Finally, numerical results show that with our proposed resource allocation scheme, the overall latency can be significantly reduced by introducing FD to MEC system. Wen He 0001, Yizhe Zhang 0003, Dazhi He, Yin Xu 0001, Yunfeng Guan 0001, Wenjun Zhang 0001 |
ICC | 7 |
| 2019 | Layered-Division Multiplexing Multicell Cooperative Multicast-Broadcast BeamformingabstractIn this paper, a layered-division multiplexing (LDM) based non-orthogonal transmission framework is proposed to enhance the spectral efficiency of Multicast- Broadcast Single Frequency Network (MBSFN). In this framework, different ranges of MBSFN areas are incorporated into a two-layer LDM system, one layer is for small scale local services, the other layer is for large scale global service. To optimize the proposed transmission framework, we design a cooperative beamforming scheme and abstract it as a max-min fair (MMF) problem. We transform the problem into the difference of convex (DC) structure and design a concave-convex procedure (CCCP) based algorithm to find a local optimal of the problem. In addition, performance upper bounds and baselines are formed through semidefinite relaxation (SDR). The results show that the proposed CCCP-based algorithm performs close to upper bounds and better than the SDR-based approach. And this LDM-based non-orthogonal transmission framework also acquires better spectral efficiency than the orthogonal transmission frameworks. Dazhi He, Yin Xu 0001, Yijia Feng, Yiwei Zhang 0015, Wenjun Zhang 0001 |
VTC Fall | 6 |
| 2019 | Beam Design for Beam Training Based Millimeter Wave V2I CommunicationsabstractIn order to achieve high-quality entertainment services and large-capacity sensor sharing, it is imperative to improve the throughput of V2I communication systems. Millimeter wave communication is a promising technology, which is generally combined with beamforming techniques. When it comes to beam design, we need to focus on the tradeoff between system throughput and alignment overhead. However, conventional optimization schemes are limited by uniform beamwidth design. Considering the characteristics of V2I communication on the highway, this paper proposes a non-uniform beamwidth design idea. Firstly, an average throughput model based on beam training is established. Then, a recursive algorithm is used to achieve non-uniform beamwidth optimization. Finally, the simulation results prove that the non-uniform beamwidth scheme can significantly improve the throughput of the V2I system. Yijia Feng, Dazhi He, Yin Xu 0001, Hongjiang Zheng, Wenjun Zhang 0001 |
VTC Fall | 6 |
| 2019 | Dynamic Stackelberg Game for Service Auction of TV White Space in 5GabstractMultimedia Broadcast Multicast Service (MBMS) will be expected to be added into 5G system in the coming 5G release of 3GPP. TV White Space (TVWS) are expected to be effectively used in MBMS mode to solve the problem of spectrum shortage. TV White Space (TVWS) can be organized to provide great help by auction to 5G service providers (SPs) in various scenarios, e.g. mobile wireless communication, Internet of Things (IOT) and Vehicular Network. We investigate imperfect information dynamic Stackelberg Game to allocate the idle TVWS spectrum resources using auction scheme, and, accordingly, the service transaction platform is constructed. First, Hidden Markov Model (HMM) is used to predict the service volume required by Service Providers (SPs). Then, the Nash Equilibrium is achieved by Service Providers, who make strategies based on the forecasting results. Next, Broadcast Operator (BO) allocates the idle spectrum by auction. The simulation results show that the predicted prices are close to the actual transaction prices and the profits of broadcasting operators and services providers are increased simultaneously. This provides solution for the unified regulation of TVWS, incenting the usage of TVWS and providing the feasible scheme that broadcasting can provide the services in 5G. Yizhe Zhang 0003, Yin Xu 0001, Dazhi He, Wen He 0001, Yunfeng Guan 0001, Wenjun Zhang 0001 |
VTC Fall | 7 |
| 2019 | An advanced bispectrum features for EEG-based motor imagery classification
Beichen Wang, Wenjun Zhang 0001 |
Expert Syst. Appl. | 5 |
| 2019 | Recognition oriented facial image quality assessment via deep convolutional neural network
Ning Zhuang, Cenhui Pan, Bingbing Ni, Yi Xu 0001, Xiaokang Yang 0001, Wenjun Zhang 0001 |
Neurocomputing | 7 |
| 2019 | Long term activity prediction in first person viewpoint
Ning Zhuang, Zefan Li, Bingbing Ni, Wenjun Zhang 0001 |
Pattern Recognit. Lett. | 6 |
| 2019 | Quality Evaluation of Image Dehazing Methods Using Synthetic Hazy ImagesabstractTo enhance the visibility and usability of images captured in hazy conditions, many image dehazing algorithms (DHAs) have been proposed. With so many image DHAs, there is a need to evaluate and compare these DHAs. Due to the lack of the reference haze-free images, DHAs are generally evaluated qualitatively using real hazy images. But it is possible to perform quantitative evaluation using synthetic hazy images since the reference haze-free images are available and full-reference (FR) image quality assessment (IQA) measures can be utilized. In this paper, we follow this strategy and study DHA evaluation using synthetic hazy images systematically. We first build a synthetic haze removing quality (SHRQ) database. It consists of two subsets: regular and aerial image subsets, which include 360 and 240 dehazed images created from 45 and 30 synthetic hazy images using 8 DHAs, respectively. Since aerial imaging is an important application area of dehazing, we create an aerial image subset specifically. We then carry out subjective quality evaluation study on these two subsets. We observe that taking DHA evaluation as an exact FR IQA process is questionable, and the state-of-the-art FR IQA measures are not effective for DHA evaluation. Thus, we propose a DHA quality evaluation method by integrating some dehazing-relevant features, including image structure recovering, color rendition, and over-enhancement of low-contrast areas. The proposed method works for both types of images, but we further improve it for aerial images by incorporating its specific characteristics. Experimental results on two subsets of the SHRQ database validate the effectiveness of the proposed measures. Xiongkuo Min, Guangtao Zhai, Ke Gu 0001, Yucheng Zhu, Jiantao Zhou 0001, Guodong Guo, Xiaokang Yang 0001, Xin-Ping Guan, Wenjun Zhang 0001 |
IEEE Trans. Multim. | 9 |
| 2019 | T-Gaming: A Cost-Efficient Cloud Gaming System at ScaleabstractCloud gaming (CG) system could pursue both high-quality gaming experience via intensive computing, and ultimate convenience anywhere at anytime through any energy-constrained mobile devices. Despite the abundance of efforts devoted, state-of-the-art CG systems still suffer from multiple key limitations: expensive deployment cost, high bandwidth consumption and unsatisfied quality of experience (QoE). As a result, existing works are not widely adopted in reality. This paper proposes a Transparent Gaming framework called T-Gaming that allows users to play any popular high-end desktop/console games on-the-fly over the Internet. T-Gaming utilizes the off-the-shelf consumer GPUs without resorting to the expensive proprietary GPU virtualization (vGPU) technology to reduce the deployment cost. Moreover, it enables prioritized video encoding based on the human visual feature to reduce the bandwidth consumption without noticeable visual quality degradation. Last but not least, T-Gaming adopts adaptive real-time streaming based on deep reinforcement learning (RL) to improve user's QoE. To evaluate the performance of T-Gaming, we implement and test a prototype system in the real world. Compared with the existing cloud gaming systems, T-Gaming not only reduces the expense per user by 75 percent hardware cost reduction and 14.3 percent network cost reduction, but also improves the normalized average QoE by 3.6-27.9 percent. Hao Chen 0036, Xu Zhang 0006, Yiling Xu, Ju Ren 0001, Jingtao Fan, Zhan Ma 0001, Wenjun Zhang 0001 |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2018 | Towards Locally Consistent Object Counting with Constrained Multi-stage Convolutional Neural Networks
Muming Zhao, Jian Zhang 0002, Wenjun Zhang 0001 |
ACCV (6) | 4 |
| 2018 | Multi-scale Spatially-Asymmetric Recalibration for Image Classification
Yan Wang 0033, Lingxi Xie, Siyuan Qiao, Ya Zhang 0002, Wenjun Zhang 0001, Alan L. Yuille |
ECCV (13) | 5 |
| 2018 | Online Multi-Object Tracking with Dual Matching Attention Networks
Ji Zhu 0002, Hua Yang 0001, Nian Liu 0002, Wenjun Zhang 0001, Ming-Hsuan Yang 0001 |
ECCV (5) | 5 |
| 2018 | Learning an Inverse Tone Mapping Network with a Generative Adversarial RegularizerabstractTransferring a low-dynamic-range (LDR) image to a high-dynamic-range (HDR) image, which is the so-called inverse tone mapping (iTM), is an important imaging technique to improve visual effects of imaging devices. In this paper, we propose a novel deep learning-based iTM method, which learns an inverse tone mapping network with a generative adversarial regularizer. In the framework of alternating optimization, we learn a U-Net-based HDR image generator to transfer input LDR images to HDR ones, and a simple CNN-based discriminator to classify the real HDR images and the generated ones. Specifically, when learning the generator we consider the content-related loss and the generative adversarial regularizer jointly to improve the stability and the robustness of the generated HDR images. Using the learned generator as the proposed inverse tone mapping network, we achieve superior iTM results to the state-of-the-art methods consistently. Shiyu Ning, Hongteng Xu, Li Song 0001, Rong Xie 0004, Wenjun Zhang 0001 |
ICASSP | 5 |
| 2018 | Modeling the Perceptual Impact of Viewport Adaptation for Immersive VideoabstractImmersive video offers the freedom to navigate inside the virtualized environment. Instead of streaming the entire bulky content, a viewport or field of view (FoV) adaptive streaming is preferred. We often stream the high-quality content within current viewport, but degraded-quality representation elsewhere, so as to reduce the network bandwidth consumption. We then could refine the quality when focusing to a new FoV. Therefore, in this work, we have attempted to model the perceptual response of the quality variations (through adapting the quantization and spatial resolution) with respect to the refinement duration, and reach at a product of two closed-form exponential functions that well explain the joint quantization and resolution induced quality impact. Analytical model is also cross-validated using another set of data with both Pearson and Spearman's rank rank correlations over 0.98. Our work would be devised to guide the bandwidth-quality optimized immersive video streaming. Shaowei Xie, Yiling Xu, Qiaojian Qian, Qiu Shen, Zhan Ma 0001, Wenjun Zhang 0001 |
ISCAS | 6 |
| 2018 | System Design of ATSC3.0 Broadcast Gateway Based on CPU-FPGAabstractAccording to the latest ATSC3.0 standard, a broadcast gateway software implementation scheme using multi-thread network programming is devised. This scheme meets real-time processing demand of fundamental system capacity. In order to satisfy extension business demand of higher capacity and higher concurrency of future fusion network, a CPU-FPGA software-hardware co-design scheme is proposed in the consideration of data processing features of broadcast gateway. Based on the results of detailed data processing task analysis, the most time consuming and the highest CPU occupation ratio tasks are distributed to FPGA implementation. Data exchange and operation synchronization between software and hardware is realized through data sharing and memory mapping I/O. According to testing results, the time consumption of main modules and the overall CPU occupation ratio are reduced effectively. The data capacity of a broadcast gateway is greatly improved at the same time. Jianhao Ding, Shuai Xiong, Dazhi He, Wenjun Zhang 0001 |
VTC Fall | 5 |
| 2018 | QoE based SDN heterogeneous LTE and WLAN multi-radio networks for multi-user accessabstractThe scarcity of the long term evolution (LTE) bandwidth and the increasing data demand call for more wireless local area networks (WLANs) to help the cellular offloading. A software-defined networking (SDN) based control approach is proposed to effectively utilize the heterogeneous LTE and WLAN radio bandwidth by operating the multi-radio interfaces simultaneously. The use of the deployed Wi-Fi access points (APs) for Internet access has been a common practice for most people, however, the total utility or the quality of experience (QoE) drops when users compete for a single Wi-Fi AP or there is no controller to aid them to allocate rates to each radio access technology (RAT). This problem will even aggravate when more APs are involved in wireless mobile networks. This paper investigates how heterogeneous resources should be coordinated and allocated to multi-user access of the LTE and Wi-Fi aggregation (LWA) network. The proposed application layer scheme can adjust the target users and resources adaptively based on estimated users states so that the total QoE attained by all users is maximized. The developed scheme can help the operator decide the associations of users to appropriate Wi-Fi APs and adaptively adjust rate allocations of each user on the associated Wi-Fi and LTE. The good performance and low complexity of the proposed scheme is validated in a network simulator (NS-3). Wei Huang 0012, De Meng, Jenq-Neng Hwang, Jounsup Park, Yiling Xu, Wenjun Zhang 0001 |
WCNC | 6 |
| 2017 | Performance Guaranteed Network Acceleration via High-Order Residual QuantizationabstractInput binarization has shown to be an effective way for network acceleration. However, previous binarization scheme could be regarded as simple pixel-wise thresholding operations (i.e., order-one approximation) and suffers a big accuracy loss. In this paper, we propose a high-order binarization scheme, which achieves more accurate approximation while still possesses the advantage of binary operation. In particular, the proposed scheme recursively performs residual quantization and yields a series of binary input images with decreasing magnitude scales. Accordingly, we propose high-order binary filtering and gradient propagation operations for both forward and backward computations. Theoretical analysis shows approximation error guarantee property of proposed method. Extensive experimental results demonstrate that the proposed scheme yields great recognition accuracy while being accelerated. Zefan Li, Bingbing Ni, Wenjun Zhang 0001, Xiaokang Yang 0001, Wen Gao 0001 |
ICCV | 3 |
| 2017 | SORT: Second-Order Response Transform for Visual RecognitionabstractIn this paper, we reveal the importance and benefits of introducing second-order operations into deep neural networks. We propose a novel approach named Second-Order Response Transform (SORT), which appends element-wise product transform to the linear sum of a two-branch network module. A direct advantage of SORT is to facilitate cross-branch response propagation, so that each branch can update its weights based on the current status of the other branch. Moreover, SORT augments the family of transform operations and increases the nonlinearity of the network, making it possible to learn flexible functions to fit the complicated distribution of feature space. SORT can be applied to a wide range of network architectures, including a branched variant of a chain-styled network and a residual network, with very light-weighted modifications. We observe consistent accuracy gain on both small (CIFAR10, CIFAR100 and SVHN) and big (ILSVRC2012) datasets. In addition, SORT is very efficient, as the extra computation overhead is less than 5%. Yan Wang 0033, Lingxi Xie, Chenxi Liu 0001, Siyuan Qiao, Ya Zhang 0002, Wenjun Zhang 0001, Qi Tian 0001, Alan L. Yuille |
ICCV | 6 |
| 2017 | Learning Mixtures of Markov Chains from Aggregate Data with Structural Constraints (Extended Abstract)abstractIn this work, we explore the learning task of mixtures of Markov chains (MMCs) from aggregate data. Our work demonstrates that although this challenging task is generally intractable because of the identifiability problem, it can be solved approximately by imposing structural constraints on its transition matrices Specifically, the proposed structural constraints include specifying active state sets corresponding to the chains and adding a series of pairwise sparse regularizers on transition matrices. Based on these two structural constraints, we propose a constrained least-squares method to learn mixtures of Markov chains. We develop a novel iterative algorithm that decomposes the overall problem into a set of convex subproblems and solves each subproblem efficiently. Experimental results on synthetic data prove that our learning method converges well and is robust to the noise in data. Moreover, the comparison with state-of-art competitors on real-world data further validates the superiority of our method. Dixin Luo, Hongteng Xu, Yi Zhen, Bistra Dilkina, Hongyuan Zha, Xiaokang Yang 0001, Wenjun Zhang 0001 |
ICDE | 7 |
| 2017 | CNN based post-processing to improve HEVCabstractIn this paper, we propose a frame-based dynamic metadata post-processing scheme in HEVC. Video sequence is classified into different categories contains complexity of video content and quality indicator for each frame, an up-to-one byte flag embedded in the bitstream is transferred as side information. Meanwhile dynamic metadata contains classification information indicates the offline training of separate network models. Specifically, we adopt a 20-layers CNN (Con-volutional Neural networks) model to extract more meaningful information from the reconstructed error and improve the filtering performance. Experimental results shows that our proposed post-processing scheme leads on average 1.6% BD-rate reduction compared with HEVC baseline on the six sequences given in 2017 ICIP Grand Challenge. Chen Li 0021, Li Song 0001, Rong Xie 0004, Wenjun Zhang 0001 |
ICIP | 4 |
| 2017 | Learning a perspective-embedded deconvolution network for crowd countingabstractWe present a novel deep learning framework for crowd counting by learning a perspective-embedded deconvolution network. Perspective is an inherent property of most surveillance scenes. Unlike the traditional approaches that exploit the perspective as a separate normalization, we propose to fuse the perspective into a deconvolution network, aiming to obtain a robust, accurate and consistent crowd density map. Through layer-wise fusion, we merge perspective maps at different resolutions into the deconvolution network. With the injection of perspective, our network is driven to learn to combine the underlying scene geometric constraints adaptively, thus enabling an accurate interpretation from high-level feature maps to the pixel-wise crowd density map. In addition, our network allows generating density map for arbitrary-sized input in an end-to-end fashion. The proposed method achieves competitive result on the WorldExpo2010 crowd dataset. Muming Zhao, Jian Zhang 0002, Fatih Porikli, Wenjun Zhang 0001 |
ICME | 5 |
| 2017 | Weight-based bit allocation scheme for VR videos in HEVCabstractSince VR videos shown on the head-mounted display (HMD) is omnidirectional, the average distortion of VR videos in all directions shall be calculated in spherical domain. Several metrics have been proposed to calculate the coding loss of VR videos in spherical domain, including S-PSNR, WS-PSNR, CPP-PSNR. The above metrics are all improved based on PSNR by creating a weight map. This paper aims to optimize the rate control scheme on HEVC mostly for WS-PSNR. According to the weight map of WS-PSNR, regions with more weights in plain are more important, thus more bits shall be allocated to the important regions. We use weight-based rate control scheme to realize the above thoughts. Weight-based rate control scheme defines bit per weight (bpw) instead of bpp. Large values of bpw indicate that the important regions deserve high bitrates, thus probably achieving better quality. Consequently, the proposed rate control scheme improves the video quality of VR videos, which leads to average gain of 2.1%, 4.3% and 1.5% of S-PSNR, WS-PSNR and CPP-PSNR. BiJia Li, Li Song 0001, Rong Xie 0004, Wenjun Zhang 0001 |
VCIP | 4 |
| 2017 | A generic method to improve no-reference image blur metric accuracy in video contentsabstractWe present in this work a generic and effective method to increase the prediction accuracy of no-reference image/video blur assessment facing the real-world content diversity. We demonstrate that benchmarking no reference image blur metrics, fitting a single logistic function to map the objective predictions to subjective scores in the well-known databases like LIVE or TID2008/2013, introduce biased fitting results towards better predictions only in the central part of the score scale. We find out that a multi-fitting approach, using the correlation parameters between subjective scores and objective predictions for content clustering and then conducting logistic fitting for each content type, can evidently improve the metric prediction accuracy in the full score scale. Besides, the overall prediction variance is also reduced with the proposed scheme, presenting more consistent results insensitive of content variation. We prove that the proposed method is of practical meaning to facilitate blur assessment techniques validated on limited databases to the vastly abundant real-life content types. Yankai Liu, Li Song 0001, Rong Xie 0004, Wenjun Zhang 0001 |
VCIP | 4 |
| 2017 | No-Reference Quality Metric of Contrast-Distorted Images Based on Information MaximizationabstractThe general purpose of seeing a picture is to attain information as much as possible. With it, we in this paper devise a new no-reference/blind metric for image quality assessment (IQA) of contrast distortion. For local details, we first roughly remove predicted regions in an image since unpredicted remains are of much information. We then compute entropy of particular unpredicted areas of maximum information via visual saliency. From global perspective, we compare the image histogram with the uniformly distributed histogram of maximum information via the symmetric Kullback-Leibler divergence. The proposed blind IQA method generates an overall quality estimation of a contrast-distorted image by properly combining local and global considerations. Thorough experiments on five databases/subsets demonstrate the superiority of our training-free blind technique over state-of-the-art full- and no-reference IQA methods. Furthermore, the proposed model is also applied to amend the performance of general-purpose blind quality metrics to a sizable margin. Ke Gu 0001, Weisi Lin, Guangtao Zhai, Xiaokang Yang 0001, Wenjun Zhang 0001, Chang Wen Chen |
IEEE Trans. Cybern. | 5 |
| 2017 | Cost-Effective Online Trending Topic Detection and Popularity Prediction in MicrobloggingabstractIdentifying topic trends on microblogging services such as Twitter and estimating those topics’ future popularity have great academic and business value, especially when the operations can be done in real time. For any third party, however, capturing and processing such huge volumes of real-time data in microblogs are almost infeasible tasks, as there always exist API (Application Program Interface) request limits, monitoring and computing budgets, as well as timeliness requirements. To deal with these challenges, we propose a cost-effective system framework with algorithms that can automatically select a subset of representative users in microblogging networks in offline, under given cost constraints. Then the proposed system can online monitor and utilize only these selected users’ real-time microposts to detect the overall trending topics and predict their future popularity among the whole microblogging network. Therefore, our proposed system framework is practical for real-time usage as it avoids the high cost in capturing and processing full real-time data, while not compromising detection and prediction performance under given cost constraints. Experiments with real microblogs dataset show that by tracking only 500 users out of 0.6 million users and processing no more than 30,000 microposts daily, about 92% trending topics could be detected and predicted by the proposed system and, on average, more than 10 hours earlier than they appear in official trends lists. Zhongchen Miao, Kai Chen 0006, Yi Fang 0008, Jianhua He 0001, Yi Zhou 0003, Wenjun Zhang 0001, Hongyuan Zha |
ACM Trans. Inf. Syst. | 6 |
| 2016 | A new AL-FEC coding scheme with limited feedbackabstractFor the next generation mobile video broadcasting, especially in-band solutions that serves the mobile devices, a limited feedback scheme via cellular channel polling is feasible to give accurate real-time information on the broadcast receivers' channel erasure rate, and decoding buffer status. In this work, we propose an AL-FEC coding degree scheme based on this feedback, to achieve a better decode efficiency and save the code redundancy. Simulation results demonstrate the effectiveness of this solution, and open up new opportunities in the next generation broadcasting system design. Wei Huang 0012, Hao Chen 0036, Yiling Xu, Zhu Li 0001, Wenjun Zhang 0001 |
MMSP | 5 |
| 2016 | Improved intra angular prediction with novel interpolation filter and boundary filterabstractIn this paper, two improved intra angular prediction methods are proposed to enhance coding performance. The first method applies new four-tap interpolation filter algorithm. The reference samples at fractional position are interpolated by DCT-based or Gaussian interpolation filter. The second method proposes extended boundary prediction filter to reduce the prediction error. The experimental results show that for AI configuration, the overall coding gain is about 0.85% on average comparing to HEVC reference software while maintaining almost the same coding time. Rujun Wei, Rong Xie 0004, Li Song 0001, Liang Zhang 0026, Wenjun Zhang 0001 |
PCS | 5 |
| 2016 | Learning a blind quality evaluation engine of screen content images
Ke Gu 0001, Guangtao Zhai, Weisi Lin, Xiaokang Yang 0001, Wenjun Zhang 0001 |
Neurocomputing | 5 |
| 2016 | A Reconfigurable Tangram Model for Scene Representation and CategorizationabstractThis paper presents a hierarchical and compositional scene layout (i.e., spatial configuration) representation and a method of learning reconfigurable model for scene categorization. Three types of shape primitives (i.e., triangle, parallelogram, and trapezoid), called tans, are used to tile scene image lattice in a hierarchical and compositional way, and a directed acyclic AND-OR graph (AOG) is proposed to organize the overcomplete dictionary of tan instances placed in image lattice, exploring a very large number of scene layouts. With certain off-the-shelf appearance features used for grounding terminal-nodes (i.e., tan instances) in the AOG, a scene layout is represented by the globally optimal parse tree learned via a dynamic programming algorithm from the AOG, which we call tangram model. Then, a scene category is represented by a mixture of tangram models discovered with an exemplar-based clustering method. On basis of the tangram model, we address scene categorization in two aspects: 1) building a tangram bank representation for linear classifiers, which utilizes a collection of tangram models learned from all categories and 2) building a tangram matching kernel for kernel-based classification, which accounts for all hidden spatial configurations in the AOG. In experiments, our methods are evaluated on three scene data sets for both the configuration-level and semantic-level scene categorization, and outperform the spatial pyramid model consistently. Tianfu Wu 0001, Song-Chun Zhu, Xiaokang Yang 0001, Wenjun Zhang 0001 |
IEEE Trans. Image Process. | 5 |
| 2016 | Learning Mixtures of Markov Chains from Aggregate Data with Structural ConstraintsabstractStatistical models based on Markov chains, especially mixtures of Markov chains, have recently been studied and demonstrated to be effective in various data mining applications such as tourist flow analysis, animal migration modeling, and transportation administration. Nevertheless, the research so far has mainly focused on analyzing data at individual levels. Due to security and privacy reasons, however, the observations in practice usually consist of coarse-grained statistics of individual data,a.k.a.aggregate data, rendering learning mixtures of Markov chains an even more challenging problem. In this work, we show that this challenging problem, although intractable in its original form, can be solved approximately by posing structural constraints on the transition matrices. The proposed structural constraints include specifying active state sets corresponding to the chains and adding a pairwise sparse regularization term on transition matrices. Based on these two structural constraints, we propose a constrained least-squares method to learn mixtures of Markov chains. We further develop a novel iterative algorithm that decomposes the overall problem into a set of convex subproblems and solves each subproblem efficiently, making it possible to effectively learn mixtures of Markov chains from aggregate data. We propose a framework for generating synthetic data and analyze the complexity of our algorithm. Additionally, the empirical results of the convergence and the robustness of our algorithm are also presented. These results demonstrate the effectiveness and efficiency of the proposed algorithm, comparing with traditional methods. Experimental results on real-world data sets further validate that our algorithm can be used to solve practical problems. Dixin Luo, Hongteng Xu, Yi Zhen, Bistra Dilkina, Hongyuan Zha, Xiaokang Yang 0001, Wenjun Zhang 0001 |
IEEE Trans. Knowl. Data Eng. | 7 |
| 2016 | Saliency-Guided Quality Assessment of Screen Content ImagesabstractWith the widespread adoption of multidevice communication, such as telecommuting, screen content images (SCIs) have become more closely and frequently related to our daily lives. For SCIs, the tasks of accurate visual quality assessment, high-efficiency compression, and suitable contrast enhancement have thus currently attracted increased attention. In particular, the quality evaluation of SCIs is important due to its good ability for instruction and optimization in various processing systems. Hence, in this paper, we develop a new objective metric for research on perceptual quality assessment of distorted SCIs. Compared to the classical MSE, our method, which mainly relies on simple convolution operators, first highlights the degradations in structures caused by different types of distortions and then detects salient areas where the distortions usually attract more attention. A comparison of our algorithm with the most popular and state-of-the-art quality measures is performed on two new SCI databases (SIQAD and SCD). Extensive results are provided to verify the superiority and efficiency of the proposed IQA technique. Ke Gu 0001, Shiqi Wang 0001, Huan Yang 0001, Weisi Lin, Guangtao Zhai, Xiaokang Yang 0001, Wenjun Zhang 0001 |
IEEE Trans. Multim. | 7 |
| 2016 | Blind Quality Assessment of Tone-Mapped Images Via Analysis of Information, Naturalness, and StructureabstractHigh dynamic range (HDR) imaging techniques have been working constantly, actively, and validly in the fault detection and disease diagnosis in the astronomical and medical fields, and currently they have also gained much more attention from digital image processing and computer vision communities. While HDR imaging devices are starting to have friendly prices, HDR display devices are still out of reach of typical consumers. Due to the limited availability of HDR display devices, in most cases tone mapping operators (TMOs) are used to convert HDR images to standard low dynamic range (LDR) images for visualization. But existing TMOs cannot work effectively for all kinds of HDR images, with their performance largely depending on brightness, contrast, and structure properties of a scene. To accurately measure and compare the performance of distinct TMOs, in this paper develop an effective and efficient no-reference objective quality metric which can automatically assess LDR images created by different TMOs without access to the original HDR images. Our model is shown to be statistically superior to recent full- and no-reference quality measures on the existing tone-mapped image database and a new relevant database built in this work. Ke Gu 0001, Shiqi Wang 0001, Guangtao Zhai, Siwei Ma 0001, Xiaokang Yang 0001, Weisi Lin, Wenjun Zhang 0001, Wen Gao 0001 |
IEEE Trans. Multim. | 7 |
| 2016 | Network Connectivity With Inhomogeneous Correlated MobilityabstractIn this paper, we derive the critical transmission range, i.e., the smallest transmission distance of nodes such that wireless network can be connected, in large-scale clustered wireless networks. Contrary to most previous literature on independent and homogeneous mobility of nodes, we consider general settings with inhomogeneous node distribution and correlated mobility. In particular, we consider three network states based on the degree of correlation among nodes, i.e., cluster-sparse state (strong correlations), cluster-dense state (weak correlations), and cluster-transitional state (medium correlations). Under each state, we focus on the following problems: 1) how to place cluster-head nodes to minimize the critical transmission range and 2) what is the corresponding minimum critical transmission range. We derive the optimal distribution of cluster-head nodes that minimizes the critical transmission range, and show that the inhomogeneous distribution of mobile nodes leads to a smaller critical transmission range. Xiaoying Liu 0001, Jinbei Zhang, Liang Liu 0013, Weijie Wu, Xiaohua Tian, Xinbing Wang, Wenjun Zhang 0001, Jun (Jim) Xu |
IEEE Trans. Wirel. Commun. | 7 |
| 2015 | Online trendy topics detection in microblogs with selective user monitoring under cost constraintsabstractAs microblog services such as Twitter become a fast and convenient communication approach, identification of trendy topics in microblog services has great academic and business value. However detecting trendy topics is very challenging due to huge number of users and short-text posts in microblog diffusion networks. In this paper we introduce a trendy topics detection system under computation and communication resource constraints. In stark contrast to retrieving and processing the whole microblog contents, we develop an idea of selecting a small set of microblog users and processing their posts to achieve an overall acceptable trendy topic coverage, without exceeding resource budget for detection. We formulate the selection operation of these subset users as mixed-integer optimization problems, and develop heuristic algorithms to compute their approximate solutions. The proposed system is evaluated with real-time test data retrieved from Sina Weibo, the dominant microblog service provider in China. It's shown that by monitoring 500 out of 1.6 million microblog users and tracking their microposts (about 15,000 daily) with our system, nearly 65% trendy topics can be detected, while on average 5 hours earlier before they appear in Sina Weibo official trends. Zhongchen Miao, Kai Chen 0006, Yi Zhou 0003, Hongyuan Zha, Jianhua He 0001, Xiaokang Yang 0001, Wenjun Zhang 0001 |
ICC | 7 |
| 2015 | Multi-Task Multi-Dimensional Hawkes Processes for Modeling Event Sequences
Dixin Luo, Hongteng Xu, Yi Zhen, Xia Ning, Hongyuan Zha, Xiaokang Yang 0001, Wenjun Zhang 0001 |
IJCAI | 7 |
| 2015 | Decorrelation-stretch based cloud detection for total sky imagesabstractCloud detection plays an important role in total-sky images based solar forecasting and has received more attention in recent years. Accurate cloud detection for complicated total-sky images is especially changeling due to the low contrast and vague boundaries between cloud and sky regions. Unlike the existing cloud detection method without any preprocessing, one novel decorrelation-stretch (DS) based method is proposed in this work, where the total-sky images are preprocessed using the DS algorithm firstly. With this enhancement, color feature disparity of cloud and sky can be intensified notably, and then a more accurate threshold can be obtained by applying the Minimum Cross Entropy (MCE) to the preprocessed image. Experimental results demonstrated the proposed scheme achieves better performance than the existing cloud detection methods on total-sky images, especially for images with low contrast or vague boundaries between cloud and sky regions. Muming Zhao, Wenjun Zhang 0001, Wei Li 0037, Jian Zhang 0002 |
VCIP | 3 |
| 2015 | An Optimized Pixel-Wise Weighting Approach for Patch-Based Image DenoisingabstractMost existing patch-based image denoising algorithms filter overlapping image patches and aggregate multiple estimates for the same pixel via weighting. Current weighting approaches always assume the restored estimates as independent random variables, which is inconsistent with the reality. In this letter, we analyze the correlation among the estimates and propose a bias-variance model to estimate the Mean Squared Error (MSE) under various weights. The new model exploits the overlapping information of the patches; it then utilizes the optimization to try to minimize the estimated MSE. Under this model, we propose a new weighting approach based on Quadratic Programming (QP), which can be embedded into various denoising algorithms. Experimental results show that the Peak Signal to Noise Ratio (PSNR) of algorithms like K-SVD and EPLL can be improved by around 0.1 dB under a range of noise levels. This improvement is promising, since it is gained independent to which image model is used, especially when the gain from designing new image models becomes less and less. Jianzhou Feng 0001, Li Song 0001, Xiaoming Huo, Xiaokang Yang 0001, Wenjun Zhang 0001 |
IEEE Signal Process. Lett. | 5 |
| 2015 | Visual Saliency Detection With Free Energy TheoryabstractVisual saliency can be thought of as the product of human brain activity. Most existing models were built upon local features or global features or both. Lately, a so-called free energy principle unifies several brain theories within one framework, and tells where easily surprise human viewers in a visual stimulus through a psychological measure. We believe that this “surprise” should be highly related to visual saliency, and thereby introduce a novel computational Free Energy inspired Saliency detection technique (FES). Our method computes the local entropy of the gap between an input image signal and its predicted counterpart that is reconstructed from the input one with a semi-parametric model. Experimental results prove that our algorithm predicts human fixation points accurately and is superior to classical/state-of-the-art competitors. Ke Gu 0001, Guangtao Zhai, Weisi Lin, Xiaokang Yang 0001, Wenjun Zhang 0001 |
IEEE Signal Process. Lett. | 5 |
| 2015 | Automatic Contrast Enhancement Technology With Saliency PreservationabstractIn this paper, we investigate the problem of image contrast enhancement. Most existing relevant technologies often suffer from the drawback of excessive enhancement, thereby introducing noise/artifacts and changing visual attention regions. One frequently used solution is manual parameter tuning, which is, however, impractical for most applications since it is labor intensive and time consuming. In this research, we find that saliency preservation can help produce appropriately enhanced images, i.e., improved contrast without annoying artifacts. We therefore design an automatic contrast enhancement technology with a complete histogram modification framework and an automatic parameter selector. This framework combines the original image, its histogram equalized product, and its visually pleasing version created by a sigmoid transfer function that was developed in our recent work. Then, a visual quality judging criterion is developed based on the concept of saliency preservation, which assists the automatic parameters selection, and finally properly enhanced image can be generated accordingly. We test the proposed scheme on Kodak and Video Quality Experts Group databases, and compare with the classical histogram equalization technique and its variations as well as state-of-the-art contrast enhancement approaches. The experimental results demonstrate that our technique has superior saliency preservation ability and outstanding enhancement effect. Ke Gu 0001, Guangtao Zhai, Xiaokang Yang 0001, Wenjun Zhang 0001, Chang Wen Chen |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2015 | No-Reference Image Sharpness Assessment in Autoregressive Parameter SpaceabstractIn this paper, we propose a new no-reference (NR)/blind sharpness metric in the autoregressive (AR) parameter space. Our model is established via the analysis of AR model parameters, first calculating the energy- and contrast-differences in the locally estimated AR coefficients in a pointwise way, and then quantifying the image sharpness with percentile pooling to predict the overall score. In addition to the luminance domain, we further consider the inevitable effect of color information on visual perception to sharpness and thereby extend the above model to the widely used YIQ color space. Validation of our technique is conducted on the subsets with blurring artifacts from four large-scale image databases (LIVE, TID2008, CSIQ, and TID2013). Experimental results confirm the superiority and efficiency of our method over existing NR algorithms, the stateof-the-art blind sharpness/blurriness estimators, and classical full-reference quality evaluators. Furthermore, the proposed metric can be also extended to stereoscopic images based on binocular rivalry, and attains remarkably high performance on LIVE3D-I and LIVE3D-II databases. Ke Gu 0001, Guangtao Zhai, Weisi Lin, Xiaokang Yang 0001, Wenjun Zhang 0001 |
IEEE Trans. Image Process. | 5 |
| 2015 | Using Free Energy Principle For Blind Image Quality AssessmentabstractIn this paper we propose a new no-reference (NR) image quality assessment (IQA) metric using the recently revealed free-energy-based brain theory and classical human visual system (HVS)-inspired features. The features used can be divided into three groups. The first involves the features inspired by the free energy principle and the structural degradation model. Furthermore, the free energy theory also reveals that the HVS always tries to infer the meaningful part from the visual stimuli. In terms of this finding, we first predict an image that the HVS perceives from a distorted image based on the free energy theory, then the second group of features is composed of some HVS-inspired features (such as structural information and gradient magnitude) computed using the distorted and predicted images. The third group of features quantifies the possible losses of “naturalness” in the distorted image by fitting the generalized Gaussian distribution to mean subtracted contrast normalized coefficients. After feature extraction, our algorithm utilizes the support vector machine based regression module to derive the overall quality score. Experiments on LIVE, TID2008, CSIQ, IVC, and Toyama databases confirm the effectiveness of our introduced NR IQA metric compared to the state-of-the-art. Ke Gu 0001, Guangtao Zhai, Xiaokang Yang 0001, Wenjun Zhang 0001 |
IEEE Trans. Multim. | 4 |
| 2014 | An efficient color image quality metric with local-tuned-global modelabstractThis paper investigates the problem of full-reference (FR) image quality assessment (IQA). In general, the ideal IQA metric should be effective and efficient, yet most of existing FR IQA methods cannot reach these two targets simultaneously. Under the supposition that the human visual perception to image quality depends on salient local distortion and global quality degradation, we introduce a novel effective and efficient local-tuned-global (LTG) model induced IQA metric. Extensive experiments are conducted on five publicly available subject-rated color image quality databases, including LIVE, TID2008, CSIQ, IVC and TID2013, to evaluate and compare our algorithm with classical and state-of-the-art FR IQA approaches. The proposed LTG is shown to work fast and outperform those competing methods. Ke Gu 0001, Guangtao Zhai, Xiaokang Yang 0001, Wenjun Zhang 0001 |
ICIP | 4 |
| 2014 | Deep learning network for blind image quality assessmentabstractNowadays, blind image quality assessment (BIQA) has been intensively studied with machine learning, such as support vector machine (SVM) and k-means. Existing BIQA metrics, however, do not perform robust for various kinds of distortion types. We believe this problem is because those frequently used traditional machine learning techniques exploit shallow architectures, which only contain one single layer of nonlinear feature transformation, and thus cannot highly mimic the mechanism of human visual perception to image quality. The recent advance of deep neural network (DNN) can help to solve this problem, since the DNN is found to better capture the essential attributes of images. We in this paper therefore introduce a new Deep learning based Image Quality Index (DIQI) for blind quality assessment. Extensive studies are conducted on the new TID2013 database and confirm the effectiveness of our DIQI relative to classical full-reference and state-of-the-art reduced- and no-reference IQA approaches. Ke Gu 0001, Guangtao Zhai, Xiaokang Yang 0001, Wenjun Zhang 0001 |
ICIP | 4 |
| 2014 | Details preservation inspired blind quality metric of tone mapping methodsabstractHigh dynamic range (HDR) images are extremely meaningful, especially in the space and medical fields. For visualization of HDR images on standard low dynamic range (LDR) display devices, how to convert HDR to LDR images naturally becomes a valuable issue, which has aroused a variety of tone-mapping operators (TMOs). To compare different LDR images created by distinct TMOs, researchers have recently provided a subject-rated tone-mapped image database, and then developed a full-reference objective tone-mapped image quality index (TMQI) based on the measurement of multi-scale signal fidelity and statistical naturalness. Instead, the basic property of HDR images about details preservation is studied in this paper. With it, a natural inference is that higher-quality tone-mapped images are capable of displaying much more details. We therefore propose a blind quality metric by estimating the amount of details in images generated by darkening/brightening an original tone-mapped images. Experimental results on the above tone-mapped image database confirm that the proposed method, despite of no reference, is robust and statistically superior to the currently optimal full-reference TMQI algorithm, and remarkably outperforms state-of-the-art no-reference IQA metrics. Ke Gu 0001, Guangtao Zhai, Min Liu 0003, Xiaokang Yang 0001, Wenjun Zhang 0001 |
ISCAS | 5 |
| 2014 | Evaluation of Different Algorithms of Nonnegative Matrix Factorization in Temporal Psychovisual ModulationabstractTemporal psychovisual modulation (TPVM) is a newly proposed information display paradigm, which can be implemented by nonnegative matrix factorization (NMF) with additional upper bound constraints on the variables. In this paper, we study all the state-of-the-art algorithms in NMF, extend them to incorporate the upper bounds and discuss their potential use in TPVM. By comparing all the NMF algorithms with their extended versions, we find that: 1) the factorization error of the truncated alternating least squares algorithm always fluctuates throughout the iterations, 2) the alternating nonnegative least squares based algorithms may slow down dramatically under the upper bound constraints, and 3) the hierarchical alternating least squares (HALS) algorithm converges the fastest and its final factorization error is often the smallest among all the algorithms. Based on the experimental results of the HALS, we propose a guideline of determining the parameter setting of TPVM, that is, the number of viewers to support and the scaling factor for adjusting the light intensity of the images formed by TPVM. This paper will facilitate the applications of TPVM. Jianzhou Feng 0001, Xiaoming Huo, Li Song 0001, Xiaokang Yang 0001, Wenjun Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2013 | Action Recognition with ActonsabstractWith the improved accessibility to an exploding amount of video data and growing demands in a wide range of video analysis applications, video-based action recognition/classification becomes an increasingly important task in computer vision. In this paper, we propose a two-layer structure for action recognition to automatically exploit a mid-level ``acton'' representation. The actons are learned via a new max-margin multi-channel multiple instance learning framework. The learned actons (with no requirement for detailed manual annotations) thus observe a property of being compact, informative, discriminative, and easy to scale. This is different from the standard unsupervised (e.g. k-means) or supervised (e.g. random forests) coding strategies in action recognition. Applying the learned actons in our two-layer structure yields the state-of-the-art classification performance on Youtube and HMDB51 datasets. Baoyuan Wang, Xiaokang Yang 0001, Wenjun Zhang 0001, Zhuowen Tu |
ICCV | 4 |
| 2013 | DR-Marker: A Novel Diminishing-Reality-Based AR RegistrationabstractThis paper proposes a novel diminishing-reality-based AR registration, called DR-marker. It is used in our proposed filmmaking system ARFMS. Although it is based on marker-based registration method, the user will not see the special marker in the real environment via diminishing reality technology. Users can apply the DR-marker just like real scene element, because it takes the advantages of marker-based and marker less-based registration approaches. After marker detection, projecting model is built to replace the marker. Then the projecting model can be tracked like a CG object. The marker may be occluded by scene elements. The virtual-to-real occlusion ensures that the DR-marker looks like the real scene elements without wrong spatial relationship. And DR-marker approach is based on a simple setup. This paper implements DR-marker in real-time. It demonstrates DR-marker is a stable and low-computation AR registration method. DR-marker can be effective used in the ARFMS. The user can perform with the CG character and stage property just like real ones. Wenjun Zhang 0001 |
ICIG | 3 |
| 2013 | Image restoration via efficient Gaussian mixture model learningabstractExpected Patch Log Likelihood (EPLL) framework using Gaussian Mixture Model (GMM) prior for image restoration was recently proposed with its performance comparable to the state-of-the-art algorithms. However, EPLL uses generic prior trained from offline image patches, which may not correctly represent statistics of the current image patches. In this paper, we extend the EPLL framework to an adaptive one, named A-EPLL, which not only concerns the likelihood of restored patches, but also trains the GMM to fit for the degraded image. To efficiently estimate GMM parameters in A-EPLL framework, we improve a recent Expectation-Maximization (EM) algorithm by exploiting specific structures of GMM from image patches, like Gaussian Scale Models. Experiment results show that A-EPLL outperforms the original EPLL significantly on several image restoration problems, like inpainting, denoising and deblurring. Jianzhou Feng 0001, Li Song 0001, Xiaoming Huo, Xiaokang Yang 0001, Wenjun Zhang 0001 |
ICIP | 5 |
| 2013 | Subjective and objective quality assessment for images with contrast changeabstractIt is widely known that, for most natural images, appropriate contrast enhancement can usually lead to improved subjective quality. Despite of its importance to image processing, contrast change has largely been overlooked in the current research of image quality assessment (IQA). To fill this void, in this paper we first report a new and dedicated contrast-changed image database (CID2013). The CID2013 database is composed of four hundred contrast-changed images of fifteen original natural images and the mean opinion scores (MOSs) recorded from twenty-two inexperienced viewers. We then proposed a novel reduced-reference image quality metric for contrast-changed images (RIQMC) using entropies and order statistics of the image histograms. Experimental results on the CID2013, TID2008, and CSIQ databases demonstrate that the proposed RIQMC metric outperforms some mainstream image quality assessment methods. Ke Gu 0001, Guangtao Zhai, Xiaokang Yang 0001, Wenjun Zhang 0001, Min Liu 0003 |
ICIP | 4 |
| 2013 | No-reference image quality assessment metric by combining free energy theory and structural degradation modelabstractIn the research of image quality assessment (IQA), no-reference approaches are usually thought of as a big challenge since none of original image information is available. To tackle this problem, we propose a new no-reference image quality metric through combining two recently proposed reduced-reference IQA models, namely the free energy based distortion metric (FEDM) and the structural degradation model (SDM). In this work, it will be shown that there exists an approximate linear relationship between the original image information of the free energy feature and the structural degradation information. Based on this observation and the application of support vector machine (SVM) that is widely used in the current study of IQA, our newly developed No-reference Free energy and Structural degradation based Distortion Metric (NFSDM) is found to alleviate the dependance of original images, and has achieved remarkably well prediction accuracy, outperforming the most two full-reference IQA approaches PSNR/SSIM and several mainstream no-reference image quality metrics. Ke Gu 0001, Guangtao Zhai, Xiaokang Yang 0001, Wenjun Zhang 0001, Longfei Liang |
ICME | 4 |
| 2013 | A new reduced-reference image quality assessment using structural degradation modelabstractImage quality assessment (IQA) is an important research area in image processing. Reduced-reference (RR) IQA methods contained therein mainly aim to estimate image quality degradations with partial information about the reference image. Following the remarkable achievement of SSIM, structural information has been recognized as one key factor, and has aroused many image quality metrics so far. In this paper, we design a structural degradation model (SDM). Then, the quality score of an image is defined as a nonlinear combination, or SVM based integration, of distance between the structural degradation information of the original and distorted images. Accordingly, a new RR IQA approach using the SDM model is exploited. Experimental results on LIVE database are provided to justify the superior prediction accuracy performance of the proposed method as compared to three significant image quality metrics, PSNR, SSIM and FEDM. Ke Gu 0001, Guangtao Zhai, Xiaokang Yang 0001, Wenjun Zhang 0001 |
ISCAS | 4 |
| 2013 | Self-adaptive scale transform for IQA metricabstractRecently, an increasing number of image quality assessment (IQA) algorithms have been developed based on multi-scale methods, such as MS-SSIM, IFC, VIF and IW-PSNR/SSIM. Inspired by the achievement of multi-scale type of IQA algorithms, this paper proposes a self-adaptive scale transform based IQA approach. Using image size and viewing distance as input variables, we construct a self-adaptive scale transform function to estimate the suitable scale transform coefficient for the following image quality metrics. Two of the most well-known full-reference IQA methods (PSNR and SSIM), and three publicly-available subjectrated image databases (LIVE, IVC and Toyama-MICT) with clear image size and viewing distance values are used as testing beds in this paper. Experimental results and comparative studies on different combinations of IQA methods and image databases suggest the effectiveness and the robustness of the proposed approach. Ke Gu 0001, Guangtao Zhai, Xiaokang Yang 0001, Wenjun Zhang 0001 |
ISCAS | 4 |
| 2013 | Brightness preserving video contrast enhancement using S-shaped Transfer functionabstractThis paper presents an efficient perceptual model inspired efficient video contrast enhancement algorithm. We propose a S-shaped transfer function for image pixel values that effectively improves the perceived contrast while preserving brightness of the scene. The S-shaped transfer function has only one control parameter that can be adaptively chosen for different video contents, such as sports, cartoon, news, and landscape programs. Then, the input image brightness is further preserved, in order to maintain the perception of human visual system (HVS) to some special scenes, such as dark scene and seaside scene. Experiments and comparative study on VQEG Phase I test database demonstrate that the proposed S-shaped Transfer function based Brightness Preserving (STBP) contrast enhancement algorithm outperforms various histogram equalization based methods such as HE, DSIHE, RSIHE and WTHE, yet with much lower computational complexity. Ke Gu 0001, Guangtao Zhai, Min Liu 0003, Xiongkuo Min, Xiaokang Yang 0001, Wenjun Zhang 0001 |
VCIP | 6 |
| 2013 | Adaptive high-frequency clipping for improved image quality assessmentabstractIt is widely known that the human visual system (HVS) applies multi-resolution analysis to the scenes we see. In fact, many of the best image quality metrics, e.g. MS-SSIM and IW-PSNR/SSIM are based on multi-scale models. However, in existing multi-scale type of image quality assessment (IQA) methods, the resolution levels are fixed. In this paper, we examine the problem of selecting optimal levels in the multi-resolution analysis to preprocess the image for perceptual quality assessment. According to the contrast sensitivity function (CSF) of the HVS, the sampling of visual information by the human eyes approximates a low-pass process. For images, the amount of information we can extract depends on the size of the image (or the object(s) inside) as well as the viewing distance. Therefore, we proposed a wavelet transform based adaptive high-frequency clipping (AHC) model to approximate the effective visual information that enters the HVS. After the high-frequency clipping, rather than processing separately on each level, we transform the filtered images back to their original resolutions for quality assessment. Extensive experimental results show that on various databases (LIVE, IVC, and Toyama-MICT), performance of existing image quality algorithms (PSNR and SSIM) can be substantially improved by applying the metrics to those AHC model processed images. Ke Gu 0001, Guangtao Zhai, Min Liu 0003, Xiaokang Yang 0001, Jun Zhou 0007, Wenjun Zhang 0001 |
VCIP | 7 |
| 2012 | Image Classification by Hierarchical Spatial Pooling with Partial Least Squares AnalysisabstractRecent coding-based image classification systems generally adopt a key step of spatial pooling operation, which characterizes the statistics of patch-level local feature codes over the regions of interest (ROI), to form the image-level representation for classification. In this paper, we present a hierarchical ROI dictionary for spatial pooling, to beyond the widely used spatial pyramid in image classification literature. By utilizing the compositionality among ROIs, it captures rich spatial statistics via an efficient pooling algorithm in deep hierarchy. On this basis, we further employ partial least squares analysis to learn a more compact and discriminative image representation. The experimental results demonstrate superiority of the proposed hierarchical pooling method relative to spatial pyramid, on three benchmark datasets for image classification. Weijia Zou, Xiaokang Yang 0001, Rui Zhang 0052, Wenjun Zhang 0001 |
BMVC | 6 |
| 2012 | Exploring controllable deterministic bits for LDPC iterative decoding in WiMAX networksabstractLow-density parity-check (LDPC) codes are playing an important role in modern wireless communication systems such as WiMAX due to their Shannon limit approaching error correction performance. To lower the decoding threshold of LDPC codes, this paper develops a novel multi-layer iterative decoding scheme using deterministic bits for multimedia communication systems. These deterministic bits serve as known information in the LDPC decoding process to reduce the redundancy during data transmission. Unlike the existing work, our proposed scheme addresses the controllable deterministic bits, such as MPEG null packets, rather than the widely investigated protocol headers. Simulation results show that our proposed scheme can achieve considerable gain in WiMAX networks. Bo Rong, Yin Xu 0001, Yiyan Wu 0001, Gilles Gagnon, Bo Liu 0001, Lin Gui 0001, Wenjun Zhang 0001 |
GLOBECOM | 7 |
| 2012 | A new no-reference stereoscopic image quality assessment based on ocular dominance theory and degree of parallax
Ke Gu 0001, Guangtao Zhai, Xiaokang Yang 0001, Wenjun Zhang 0001 |
ICPR | 4 |
| 2012 | Nonlinear additive model based saliency map weighting strategy for image quality assessmentabstractMost state-of-the-art image quality metrics are based on the two-step approach: local distortion/fidelity measurement and pooling. During the pooling stage, many weighting strategies have been proposed incorporating properties of the distortion itself, various masking effects and visual attention. Recently, researchers have devoted great enthusiasm and effort to the improvement of image quality assessment using visual saliency models. In this research, it is noticed that visual saliency features of both the original image and the distorted one have impacts on the process of image quality assessment. To reduce the overlapping effects, a nonlinear additive model is proposed to integrate saliency features from the original and distorted images towards improved error weighting results. Our extensive experimental studies on four publicly available image databases (LIVE, TID2008, CSIQ and A57) indicate that the proposed improved nonlinear additive model based saliency map weighting strategy constantly leads to higher prediction accuracy for image quality assessment than traditional methods. Ke Gu 0001, Guangtao Zhai, Xiaokang Yang 0001, Li Chen 0021, Wenjun Zhang 0001 |
MMSP | 5 |
| 2012 | New bounds on image denoising: Viewpoint of sparse representation and non-local averagingabstractImage denoising plays a fundamental role in many image processing applications. Utilizing sparse representation and nonlocal averaging together is such a successful framework that leads to considerable progress in denoising. Almost all the newly proposed denoising algorithms are built base on it, different in detailed implementation, and the denoising performance seems converging. What is the denoising bound of this framework turns into a key question. In this paper, we assume all the possible algorithms under the framework can be approximated by a fixed two steps denoising process with different parameters. Step one cluster geometric similar image patches into groups so that patches within each group could be sparse represented under the basis of the group. Step two use the atoms of the group basis and radiometric similar patches of each patch for non-local averaging. The parameters of the process are the cluster number, the atoms and the number of radiometric similar patches for estimating each patch. Finally, the bound is derived as the minimum denoising error of all the possible parameters. Comparing with previous bounds, the new one is image specific and more practical. Experiment results show that there still exists room to improve the denoising performance for natural images. Jianzhou Feng 0001, Li Song 0001, Xiaoming Huo, Xiaokang Yang 0001, Wenjun Zhang 0001 |
VCIP | 5 |
| 2012 | Learning reconfigurable scene representation by tangram modelabstractThis paper proposes a method to learn reconfigurable and sparse scene representation in the joint space of spatial configuration and appearance in a principled way. We call it the tangram model, which has three properties: (1) Unlike fixed structure of the spatial pyramid widely used in the literature, we propose a compositional shape dictionary organized in an And-Or directed acyclic graph (AOG) to quantize the space of spatial configurations. (2) The shape primitives (called tans) in the dictionary can be described by using any “off-the-shelf” appearance features according to different tasks. (3) A dynamic programming (DP) algorithm is utilized to learn the globally optimal parse tree in the joint space of spatial configuration and appearance. We demonstrate the tangram model in both a generative learning formulation and a discriminative matching kernel. In experiments, we show that the tangram model is capable of capturing meaningful spatial configurations as well as appearance for various scene categories, and achieves state-of-the-art classification performance on the LSP 15-class scene dataset and the MIT 67-class indoor scene dataset. Tianfu Wu 0001, Song-Chun Zhu, Xiaokang Yang 0001, Wenjun Zhang 0001 |
WACV | 5 |
| 2012 | A Psychovisual Quality Metric in Free-Energy PrincipleabstractIn this paper, we propose a new psychovisual quality metric of images based on recent developments in brain theory and neuroscience, particularly the free-energy principle. The perception and understanding of an image is modeled as an active inference process, in which the brain tries to explain the scene using an internal generative model. The psychovisual quality is thus closely related to how accurately visual sensory data can be explained by the generative model, and the upper bound of the discrepancy between the image signal and its best internal description is given by the free energy of the cognition process. Therefore, the perceptual quality of an image can be quantified using the free energy. Constructively, we develop a reduced-reference free-energy-based distortion metric (FEDM) and a no-reference free-energy-based quality metric (NFEQM). The FEDM and the NFEQM are nearly invariant to many global systematic deviations in geometry and illumination that hardly affect visual quality, for which existing image quality metrics wrongly predict severe quality degradation. Although with very limited or even without information on the reference image, the FEDM and the NFEQM are highly competitive compared with the full-reference SSIM image quality metric on images in the popular LIVE database. Moreover, FEDM and NFEQM can measure correctly the visual quality of some model-based image processing algorithms, for which the competing metrics often contradict with viewers' opinions. Guangtao Zhai, Xiaolin Wu 0001, Xiaokang Yang 0001, Weisi Lin, Wenjun Zhang 0001 |
IEEE Trans. Image Process. | 5 |
| 2011 | Capacity Theorem and Optimal Power Allocation for Frequency Division Multiple-Access Relay ChannelsabstractThe capacity region of general Multiple-Access Relay Channel (MARC) has long been an open question. In this paper, we consider MARC with orthogonal components, where the channel from the source senders to the relay node is orthogonal to the channel from the source senders and relay to the destination. We give the capacity theorem for the discrete memoryless MARC with orthogonal components. Based on the result, we obtain the capacity region of frequency-division Gaussian MARC. The optimal power allocation for the source nodes to allocate power on different frequencies is investigated. Numerical results show that the optimal allocation of power can achieve the maximum sum rate. Junhua Liang, Xinbing Wang, Wenjun Zhang 0001 |
GLOBECOM | 3 |
| 2011 | Learning sparse dictionaries with a popularity-based modelabstractSparse signal representation based on overcomplete dictionaries has recently been extensively investigated, rendering the state-of-the-art results in signal, image and video processing. We propose a novel dictionary learning algorithm-the PK-SVD algorithm-which assumes prior probabilities on the dictionary atoms and learns a sparse dictionary under a popularity-based model. The prior distribution brings the flexibility that is desirable in applications. We examine our algorithm in both synthetic tests and image denoising experiments. Jianzhou Feng 0001, Li Song 0001, Xiaoming Huo, Xiaokang Yang 0001, Wenjun Zhang 0001 |
ICASSP | 5 |
| 2011 | Learning dictionary via subspace segmentation for sparse representationabstractSparse signal representation based on redundant dictionaries contributed to much progress in image processing in the past decades. But the common overcomplete dictionary model is not well structured and there is still no guideline for selecting the proper dictionary size. In this paper, we propose a new algorithm for dictionary learning based on subspace segmentation. Our algorithm divides the training data into sub-spaces and constructs the dictionary by extracting the shared basis from multiple subspaces. The learned dictionary is well structured and its size is adaptive to the training data. We analyze this algorithm and demonstrate its ability on some initial supportive experiments using real image data. Jianzhou Feng 0001, Li Song 0001, Xiaokang Yang 0001, Wenjun Zhang 0001 |
ICIP | 4 |
| 2011 | Example-based image contrast enhancementabstractIn this paper, a novel example-based contrast enhancement algorithm is proposed. The proposed approach enhances the contrast by learning some important informative priors from the histogram of the example image. The experimental results indicate that the proposed Example-based Dist-Stretched (ExDS) contrast enhancement algorithm can boost the image contrast effectively. And thanks to the example-based learning process, the output images from the ExDS algorithm have more natural looking than those of traditional histogram equalization based methods. The proposed ExDS algorithm can also be extended to the applications of contrast correction for old film restoration as well as tone mapping for image and video post-productions. Xiaokang Yang 0001, Li Chen 0021, Guangtao Zhai, Wenjun Zhang 0001 |
MMSP | 5 |
| 2011 | V-OFDM: On Performance Limits over Multi-Path Rayleigh Fading ChannelsabstractAs a bridge of connecting orthogonal frequency division multiplexing (OFDM) with single-carrier frequency domain equalization (SC-FDE) techniques, Vector OFDM (V-OFDM) provides significant flexibility in system design. This paper presents an analytical study of V-OFDM over multi-path fading channels. Our goal is to investigate the diversity gain and coding gain of each vector block (VB) in V-OFDM so as to ultimately reveal its performance limits over fading channel. By using algebraic number theory tools, we rigorously prove for the first time that a majority of VBs in V-OFDM can surely realize the diversity gain of min {M,G} , where M is the length of each VB, and G is the total number of channel taps. Furthermore, some specific VBs, whose length equals the total number of channel taps, can not only harvest the maximum diversity gain but also achieve the maximum coding gain. It is further demonstrated that, even though VBs fail to benefit from additional diversity gain when M exceeds G, they can enjoy significantly increased coding gains. Our analysis concludes that it is preferable to choose the length of VBs to be equal to the number of channel taps in consideration of both overall system performance and computational complexity. Peng Cheng 0002, Meixia Tao, Yue Xiao 0001, Wenjun Zhang 0001 |
IEEE Trans. Commun. | 4 |
| 2011 | Adaptive Sequential Prediction of Multidimensional Signals With Applications to Lossless Image CodingabstractWe investigate the problem of designing adaptive sequential linear predictors for the class of piecewise autoregressive multidimensional signals, and adopt an approach of minimum description length (MDL) to determine the order of the predictor and the support on which the predictor operates. The design objective is to strike a balance between the bias and variance of the prediction errors in the MDL criterion. The predictor design problem is particularly interesting and challenging for multidimensional signals (e.g., images and videos) because of the increased degree of freedom in choosing the predictor support. Our main result is a new technique of sequentializing a multidimensional signal into a sequence of nested contexts of increasing order to facilitate the MDL search for the order and the support shape of the predictor, and the sequentialization is made adaptive on a sample by sample basis. The proposed MDL-based adaptive predictor is applied to lossless image coding, and its performance is empirically established to be the best among all the results that have been published till present. Xiaolin Wu 0001, Guangtao Zhai, Xiaokang Yang 0001, Wenjun Zhang 0001 |
IEEE Trans. Image Process. | 4 |
| 2010 | Bayesian error concealment with DCT pyramidabstractIn this paper, the problem of concealing missing image/video blocks is casted into a framework of Bayesian estimation. The conditional expectation of the missing block vector is taken over a pilot vector of correctly decoded pixels near the missing block. Multiple observations of the missing vector and pilot vectors obtained in a nonlocal manner are used to approximate the expectation. We design a multiscale estimation approach with DCT pyramid to improve estimation efficiency. The DC image of the missing block is recovered first, and then more details related to high frequency AC coefficients are recovered successively. The algorithm is found to be quite competitive among state-of-the-art, and more substantial improvement over existing algorithms on image with heavy loss rate/large block size is observed in our experiments. Guangtao Zhai, Xiaokang Yang 0001, Weisi Lin, Wenjun Zhang 0001, Yi Xu 0001 |
ICASSP | 4 |
| 2010 | Simultaneous deblocking and error concealment for decoded visual signalabstractWe propose a unified deblocking/error-concealment algorithm for simultaneously alleviating blockiness and restoring lost blocks. Blockiness and block loss are two of the most frequently encountered artifacts for block-DCT based visual signal coding and transmission. The problem of error-concealment and deblocking is formulated as block-wise expectation estimation processes conditioned on available pixels in a local neighborhood. A nonparametric kernel regression approach is used for approximating the conditional probability and thus the conditional expectation of each image block. Missing blocks are restored and blockiness are suppressed with a single pass of the algorithm. The algorithm is highly efficient in that only block-wise additions and multiplications in pixel domain within a local neighborhood are involved. Experimental results are provided to justify the effectiveness of the proposed unified error-concealment/deblocking algorithm. Guangtao Zhai, Xiaokang Yang 0001, Weisi Lin, Wenjun Zhang 0001 |
ISCAS | 4 |
| 2010 | Image denoising using local tangent space alignmentabstractWe propose a novel image denoising approach, which is based on exploring an underlying (nonlinear) lowdimensional manifold. Using local tangent space alignment (LTSA), we 'learn' such a manifold, which approximates the image content effectively. The denoising is performed by minimizing a newly defined objective function, which is a sum of two terms: (a) the difference between the noisy image and the denoised image, (b) the distance from the image patch to the manifold. We extend the LTSA method from manifold learning to denoising. We introduce the local dimension concept that leads to adaptivity to different kind of image patches, e.g. flat patches having lower dimension. We also plug in a basic denoising stage to estimate the local coordinate more accurately. It is found that the proposed method is competitive: its performance surpasses the K-SVD denoising method. Jianzhou Feng 0001, Li Song 0001, Xiaoming Huo, Xiaokang Yang 0001, Wenjun Zhang 0001 |
VCIP | 5 |
| 2010 | A Modified Belief Propagation Algorithm Based on Attenuation of the Extrinsic LLRabstractIn this paper, we propose a modification to Belief Propagation (BP) decoding algorithm for LDPC codes. The modification is to attenuate the check to bit extrinsic logarithm likelihood ratio by a factor α, when sudden sign change happens. This modification can be applied to both the standard BP algorithm and the joint row and column (JRC) BP algorithm. Simulation results show that the BER and WER performance of both traditional BP and JRC BP algorithms is improved by this method. The expense of the proposed modification is a slight increase in the average number of decoding iterations. Yin Xu 0001, Bo Liu 0001, Lin Gui 0001, Bo Rong, Yiyan Wu 0001, Wenjun Zhang 0001 |
VTC Fall | 7 |
| 2010 | Designing LDPC Codes with Gated Noise Model for Terrestrial Mobile DTV ChannelsabstractThis paper investigates the design of LDPC codes over mobile DTV multipath channels. Most of the existing LDPC codes are optimized for additive white Gaussian noise (AWGN) channel, and not feasible to encounter the long burst error occurring in mobile DTV channel. Accordingly, we study the error propagation statistics of decision feedback equalizer (DFE) and formulate it into a gated noise model. To achieve good error correction in burst error channel, we proposed a class of dual-degree IRA codes to balance the metrics of decoding threshold and robustness. Extensive simulation results are presented in this paper to justify the performance of dual-degree IRA codes over gated noise model. Bo Liu 0001, Yin Xu 0001, Bo Rong, Yiyan Wu 0001, Gilles Gagnon, Lin Gui 0001, Wenjun Zhang 0001 |
VTC Fall | 8 |
| 2010 | Mobile Location Finding Using ATSC Mobile/Handheld Digital TV RF Watermark SignalsabstractThis paper investigates the use of ATSC M/H digital television (DTV) signal for location finding. In comparison to satellite based location finding system, DTV signals have higher field strength, wider bandwidth, lower frequency band, and DTV transmission towers are pervasively available everywhere. They can be used for indoor and mobile location finding in major cities where satellite based system might not function well. The ATSC receiver can obtain the multiple transmitter impulse responses and signal arrival times using the embedded RF watermark (RFWM) signal, and then derives its geographic coordinates based on the position of ATSC transmitters. As a critical step of this process, the transmitter identification in mobile environment has significant impact on the overall accuracy of location finding. In this paper, we present extensive analytical and simulation results to demonstrate the performance of RFWM technology over mobile channels. Bo Rong, Bo Liu 0001, Yiyan Wu 0001, Gilles Gagnon, Lin Gui 0001, Wenjun Zhang 0001 |
VTC Fall | 6 |
| 2010 | Joint Compressive Sensing in Wideband Cognitive NetworksabstractIn this paper, a distributed compressive spectrum sensing scheme in wideband cognitive radio networks is discussed. An AIC RF front-end sampling structure is proposed requiring only low rate ADCs and few storage units for spectrum sampling. Multiple CRs collect compressed samples through AICs and recover spectrum jointly. A novel joint sparsity model is defined in this scenario, along with a universal recovery algorithm based on S-OMP. Numerical simulations show this algorithm outperforms current existing algorithms under this model and works competently under other existing models. Junhua Liang, Wenjun Zhang 0001, Youyun Xu, Xiaoying Gan, Xinbing Wang |
WCNC | 3 |
| 2010 | Introduction of the TCSVT Associate Editors
Roberto Rinaldo, Alberto Signoroni, Raouf Hamzaoui, Wenjun Zhang 0001, Riccardo Bernardini, Francesco G. B. De Natale, Rastislav Lukac, Anthony Vetro, Houqiang Li |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2010 | Bayesian Error Concealment With DCT Pyramid for ImagesabstractIn this paper, the problem of concealing missing image blocks is casted into a framework of Bayesian estimation. The conditional expectation of the missing block vector is taken over a pilot vector of correctly decoded pixels near the missing block. Multiple observations of the missing vector and pilot vectors obtained in a neighborhood are used to approximate the expectation. We design a multiscale estimation approach with discrete cosine transform pyramid to improve estimation efficiency. The DC image of the missing block is recovered first, and then more details related to high-frequency AC coefficients are recovered successively. Moreover, the algorithm operates in an iterative mode through using estimated block to refine the searching process for the next estimation. The algorithm is found to perform very well for a wide range of block loss rates. Substantial improvement over 14 existing error concealment (inclusive of inpainting) algorithms on various images is demonstrated in our extensive experiments, under different test conditions inclusive of high-loss rates and large block sizes. Guangtao Zhai, Xiaokang Yang 0001, Weisi Lin, Wenjun Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2010 | Directional Lapped Transforms for Image CodingabstractIn this paper, we present the design of directional lapped transforms for image coding. A lapped transform, which can be implemented by a prefilter followed by a discrete cosine transform (DCT), can be factorized into elementary operators. The corresponding directional lapped transform is generated by applying each elementary operator along a given direction. The proposed directional lapped transforms are not only nonredundant and perfectly reconstructed, but they can also provide a basis along an arbitrary direction. These properties, along with the advantages of lapped transforms, make the proposed transforms appealing for image coding. A block-based directional transform scheme is also presented and integrated into HD Phtoto, one of the state-of-the-art image coding systems, to verify the effectiveness of the proposed transforms. Jizheng Xu, Feng Wu 0001, Jie Liang 0001, Wenjun Zhang 0001 |
IEEE Trans. Image Process. | 4 |
| 2009 | Incorporating TCP Acknowledgements in MAC Layer in IEEE 802.11 Multihop Ad Hoc NetworksabstractThe poor performance of TCP in multihop ad hoc networks is mainly attributed to the inefficient interaction among different protocol layers in previous literature, while the heavy load caused by end-to-end TCP acknowledgements (ACKs) with limited information is usually ignored. In this paper, we propose a novel incorporating ACK transfer scheme, IACK, to alleviate its impact. In IACK, TCP acknowledgements are incorporated in the control packets at the MAC layer and are transferred hop by hop from the sink node to the source node. To meet the requirement of IACK, we enhance the packet queuing policy at the routing layer, and propose a new rate-based TCP transfer scheme, TCP-AP+. Then, we implement IACK in ns-2, evaluate it over comprehensive scenarios and compare it with TCP-AP and TCP-Newreno. Simulation results show that IACK improves both the TCP throughput and goodput significantly. Lianghui Ding, Wenjun Zhang 0001, Hui Yu 0002, Xinbing Wang, Youyun Xu |
GLOBECOM | 2 |
| 2009 | Joint Scheduling and Relay Selection in One- and Two-Way Relay Networks with BufferingabstractIn most wireless relay networks, the source and relay nodes transmit successively via fixed time division (FTD) and each relay forwards a packet immediately upon receiving. In this paper we enable the buffering capability of relay nodes and propose a framework for joint scheduling and relay selection. The goal is to maximize the system long-term throughput by fully exploiting multi-user diversity in the network. We develop two joint scheduling and relay selection (JSRS) algorithms for unidirectional and bidirectional traffic, respectively. The novel cross-layer relay selection metrics which our algorithms are based upon take into account both instantaneous channel conditions and the queuing status. We also demonstrate that the proposed JSRS can be realized in a distributed way without explicit coordination among the network nodes. Extensive simulation is carried out to evaluate the performance of the proposed JSRS with buffering in comparison with traditional FTD without buffering. Typical throughput enhancements up to 101% and 110% are observed in one-way and two-way relay networks respectively, at low signal-to-noise ratio (0 dB). Lianghui Ding, Meixia Tao, Wenjun Zhang 0001 |
ICC | 4 |
| 2009 | Sub clustering K-SVD: Size variable dictionary learning for sparse representationsabstractSparse signal representation from overcomplete dictionaries have been extensively investigated in recent research, leading to state-of-the-art results in signal, image and video restoration. One of the most important issues is involved in selecting the proper size of dictionary. However, the related guidelines are still not established. In this paper, we tackle this problem by proposing a so-called sub clustering K-SVD algorithm. This approach incorporates the subtractive clustering method into K-SVD to retain the most important atom candidates. At the same time, the redundant atoms are removed to produce a well-trained dictionary. As for a given dataset and approximation error bound, the proposed approach can deduce the optimized size of dictionary, which is greatly compressed as compared with the one needed in the K-SVD algorithm. Jianzhou Feng 0001, Li Song 0001, Xiaokang Yang 0001, Wenjun Zhang 0001 |
ICIP | 4 |
| 2009 | On the coding gain of intra-predictive transformsabstractIntra-predictive transforms are a kind of block-based transforms that can exploit both the intra- and inter-block correlations. This paper analyzes the coding gains of intra-predictive transforms for the Gaussian process. The tight upper bound of the coding gain is derived, which is shown to be better than both the discrete cosine transform and the Karhunen Loe¿ve transform. The optimal intra-predictive transform that achieves the upper bound is also presented. Actual coding results on images verify the effectiveness of the optimal intra-predictive transform. Jizheng Xu, Feng Wu 0001, Wenjun Zhang 0001 |
ICIP | 3 |
| 2009 | MDL context modeling of images with application to denoisingabstractThe lately popularized patch-based nonlocal (NL) image processing approach is cast into a framework of statistical context modeling, a thoroughly studied topic in data compression and information theory. The adaptation of image patch (context) to local waveform is crucial to the performance of NL-type of image processing but yet lacks a rigorous study. In this paper we propose a minimum description length (MDL) approach for choosing the size and spatial configuration of the context in which a degraded pixel is to be restored. The MDL criterion of context formation aims to strike an optimal balance between the variance and bias of the errors in fitting a 2D piecewise autoregressive (PAR) model to input image signal. To exemplify the use of the proposed context modeling technique in image processing, an MDL-guided context-based image denoiser is derived and its performance evaluated. Empirical results show that the new context-based denoiser is highly competitive against the current state of the art. Guangtao Zhai, Xiaolin Wu 0001, Xiaokang Yang 0001, Wenjun Zhang 0001 |
ICIP | 4 |
| 2009 | Intra-predictive Transforms for Image CodingabstractThis paper presents a new transform framework, intra-predictive transforms, for block-based image coding. The proposed framework unifies cross-block prediction and transforms within a block. Thus, it can exploit both the intra- and inter- block correlations. The intra-prediction in H.264 can be viewed as a special case of intra-predictive transforms. The new framework provides more flexibility to design new transform for coding. As an example, a new intra-predictive transform based on frequency domain prediction is also presented to show the advantages of the proposed framework. Jizheng Xu, Feng Wu 0001, Wenjun Zhang 0001 |
ISCAS | 3 |
| 2009 | Frame Rate Up-conversion with Edge-weighted Motion Estimation and Trilateral InterpolationabstractWe propose in this paper a novel video frame-rate up-conversion method using edge-weighted motion estimation and trilateral filtering. First, the method produces two frames of the intermediate frame from its previous and following frames with estimated motion vectors. It then obtains an initial estimate from the two frames and calculates its pixel reliability. Finally, a trilateral filter is applied on the initial estimate to correct the unreliable pixels and missing pixels. Compared with the existing methods, the proposed method not only reduces the computation cost on motion vector estimation, but also suppresses the interpolation noises and motion compensation errors. Experimental results show that the proposed method provides better subjective and objective quality, and obtains up to 4-dB PSNR improvement. Lei Zhang 0006, Ci Wang, Wenjun Zhang 0001, Yap-Peng Tan |
ISCAS | 3 |
| 2009 | Efficient quadtree based block-shift filtering for deblocking and deringing
Guangtao Zhai, Weisi Lin, Jianfei Cai 0001, Xiaokang Yang 0001, Wenjun Zhang 0001 |
J. Vis. Commun. Image Represent. | 5 |
| 2009 | An improved block size selection method based on macroblock movement characteristic
Jun Sun 0005, Rong Xie 0004, Songyu Yu, Wenjun Zhang 0001 |
Multim. Tools Appl. | 5 |
| 2009 | Robust Video Region-of-Interest Coding Based on Leaky PredictionabstractA video region-of-interest (ROI) scalable coding scheme can ensure the priority of ROI. Error protection schemes can be used to guarantee the correct receipt of the ROI stream when transporting ROI scalable video over an error-prone network. However, we find that the correct receipt of ROI bitstreams cannot ensure the correct decoding of ROI due to the unique issue of the cross error propagation between ROI and background in ROI scalable coding. In this letter, we propose an ROI scalable coding framework based on leaky prediction (LP) for robustly transporting video over an error-prone network. Although several LP approaches have been proposed to improve layered coding, they cannot be applied to ROI scalable coding straightforwardly due to the cross error propagation issue. We deploy a leaky factor to weigh the two predictions: one from the constrained motion estimation (ME) within the ROI layer of the reference frame, and the other from the unrestricted ME in the overall reference frame. Simulation results show that the proposed scheme enhances the robustness of ROI scalability while maintaining coding efficiency. Qian Chen 0024, Xiaokang Yang 0001, Li Song 0001, Wenjun Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2009 | Robust Color Demosaicking With Adaptation to Varying Spectral CorrelationsabstractAlmost all existing color demosaicking algorithms for digital cameras are designed on the assumption of high correlation between red, green, blue (or some other primary color) bands. They exploit spectral correlations between the primary color bands to interpolate the missing color samples, but in areas of no or weak spectral correlations, these algorithms are prone to large interpolation errors. Such demosaicking errors are visually objectionable because they tend to correlate with object boundaries and edges. This paper proposes a remedy to the above problem that has long been overlooked in the literature. The main contribution of this work is a hybrid demosaicking approach that supplements an existing color demosaicking algorithm by combining its results with those of adaptive intraband interpolation. This is formulated as an optimal data fusion problem, and two solutions are proposed: one is based on linear minimum mean-square estimation and the other based on support vector regression. Experimental results demonstrate that the new hybrid approach is more robust and eliminates the worst type of color artifacts of existing color demosaicking methods. Fan Zhang 0080, Xiaolin Wu 0001, Xiaokang Yang 0001, Wenjun Zhang 0001, Lei Zhang 0006 |
IEEE Trans. Image Process. | 4 |
| 2009 | Channel estimation joint DOAs and time delay correlation in TD-SCDMA mobile radio systemsabstractAbstract An improved channel estimation technique based on the Steiner low‐cost channel estimator is proposed, which is widely used in TD‐SCDMA (Time Division‐Synchronous Code Division Multiple Access) cellular mobile radio systems. TD‐SCDMA is also known as third‐generation mobile systems where adaptive antennas are employed. As additive noise has a great adverse effect on the performance of the Steiner estimator, the proposed method employs time‐correlated post‐processing with a threshold filter to reduce channel noise and compensate channel variations. Furthermore, channel estimation combining direction‐of‐arrivals (DOAs) is performed, which can reduce channel interferences without adding computational complexity, for the information of DOAs has been obtained by the inherent adaptive antenna system. The performance of the improved channel estimator is compared with conventional channel estimation approaches, and numerical results show that the new approach can lead to considerable performance enhancement even in high‐speed vehicle propagation environments. Copyright © 2008 John Wiley & Sons, Ltd. Zhinian Luo, Wenjun Zhang 0001, Youyun Xu |
Wirel. Commun. Mob. Comput. | 2 |
| 2008 | Full-chip leakage analysis in nano-scale technologies: mechanisms, variation sources, and verificationabstractIn this paper, a methodology for full-chip leakage analysis based on accurate modeling of different leakage currents in nano-scaled MOSFETs has been developed. Novel process effects have been covered in our statistical model, and a systematic characterization method of leakage-related parameter variations has been proposed. With these two contributions, we present an effective algorithm to address the growing issue of full-chip leakage verification for actual-fabrication circuits. Unlike many traditional approaches that rely on log-Normal approximations, the proposed algorithm applies a quadratic model of the logarithm for the full-chip leakage current and is able to include both Gaussian and non-Gaussian parameter distributions. Our simulation examples in a 65 nm CMOS process demonstrate that the proposed methodology provides more accurate results compared with the previous methods, while achieving orders of magnitude more efficiency than a Monte Carlo analysis. Wenjun Zhang 0001, Zhiping Yu |
DAC | 2 |
| 2008 | Directional Lapped Transforms for Image CodingabstractThis paper presents a scheme to design directional lapped transforms. Lapped transforms can be factorized into lifting steps. By introducing directional operator into each lifting step, the directional lapped transform is constructed. The directional lapped transform proposed not only preserves the advantages of lapped transforms, it also can represent directional signals more efficiently. An image coding scheme using the directional lapped transform is also described. Compared to the state-of-the-art image coding using lapped transform, HD photo, the proposed scheme shows more than 20 dB's gain for artificial images with strong directional correlations. And for natural images, up to 1.5 dB's gain can also be observed. Jizheng Xu, Feng Wu 0001, Jie Liang 0001, Wenjun Zhang 0001 |
DCC | 4 |
| 2008 | Learning object classes from image thumbnails through deep neural networksabstractWe propose a new approach for recognizing object classes which is based on the intuitive idea that human beings are able to perform the task well given only thumbnails (coarse scale version) of images. Unlike previous work which uses local image features at fine scales, our approach uses thumbnails directly, and captures their high-order correlations at coarse scales through deep multi-layer neural networks based on restricted Boltzmann machines. Specifically, the pretraining stage of such networks takes on the role of feature extraction. Experimental results show that the proposed approach is comparable to other state-of-the-art recognition methods in terms of accuracy. The merits of the proposed approach come from the simplicity of the workflow and the parallelizability of the implementation structure. Erkang Chen, Xiaokang Yang 0001, Hongyuan Zha, Rui Zhang 0052, Wenjun Zhang 0001 |
ICASSP | 5 |
| 2008 | An improved DMVE temporal error concealmentabstractDuring video transmission over error-prone networks, the compressed bit stream is often corrupted by channel errors which may cause the video quality degrading suddenly. In this paper, we present a novel temporal error concealment technique as a post-processing tool at the decoder side for recovering the lost information. In order to recover the lost motion vector, an improved decoder motion vector estimation (DMVE) criterion is introduced which considers temporal correlation and motion trajectory together. We utilize pixels in the two previous frames as well as surrounding pixels of the lost block. The best motion vector is determined according to the criterion, and then the lost pixels are recovered using motion compensation. Simulations show that the proposed technique can achieve remarkable objective (PSNR) and subjective gains in the quality of the recovered video. Yang Ling, Yunqiang Liu, Wenjun Zhang 0001 |
ICASSP | 4 |
| 2008 | Scalable visual sensitivity profile estimationabstractWe propose a computational model for estimating scalable visual sensitivity profile (SVSP) of video, which is a hierarchy of saliency maps that simulates the bottom-up and top- down attention of the human visual system (HVS). The bottom- up process considers low level stimulus-driven visual features such as intensity, color, orientation and motion. The top-down process simulates the high level task-driven cognitive features such as finding human faces and captions in the video. The nonlinear addition model has been used for integrating low level visual features. A full center-surrounded receptive field profile is introduced to provide spatial scalability of the model. Due to the hierarchical nature, the proposed SVSP can be directly used to augment the visual quality of codings with spatial scalability. To justify the effectiveness of the proposed SVSP, extended experiments of its application in visual quality assessment are conducted. Guangtao Zhai, Qian Chen 0024, Xiaokang Yang 0001, Wenjun Zhang 0001 |
ICASSP | 4 |
| 2008 | Vegas-W: An Enhanced TCP-Vegas for Wireless Ad Hoc NetworksabstractThe performance of TCP-Vegas is not satisfactory in multihop ad hoc networks over IEEE 802.11 MAC protocol. We analyze the problem with a unified network model and simulation results. We observe that the aggregate throughput of all traffics decreases as the load of the network increases. The main reasons lie in Vegas's large minimum congestion window, large reset slow start threshold and aggressive window increase policy. To fix these problems, we propose a modified TCP protocol based on TCP-Vegas for multihop ad hoc networks, called Vegas- W. We extend the congestion window to fraction; change the probing mechanisms of legacy TCP-Vegas in both slow start and congestion avoidance and update slow start threshold tracking the stable window. We evaluate the performance of Vegas-W through ns-2. Extensive simulation results under a variety of scenarios show that Vegas-W can improve the throughput up to 87% over legacy TCP-Vegas and up to 27% over FeW, which is another improved algorithm based on TCP-Newreno scenarios. Lianghui Ding, Xinbing Wang, Youyun Xu, Wenjun Zhang 0001, Wen Chen 0001 |
ICC | 4 |
| 2008 | Spatial and temporal correlation based frame rate up-conversionabstractIn this paper, we present a new motion-compensated interpolation method for frame rate up-conversion based on spatial and temporal correlation. First, the motion vectors (MVs) embedded in the bit-stream are examined and corrected by using the spatial correlation in the MV field. Second, motion trajectory based MV estimation for the block in the interpolated frame is used to gain the initial value. Third, further refinement on the estimated MV is performed based on consideration of the spatial correlation in the interpolated frame and temporal correlation in the before and after frames together. Finally, bi-directional interpolation is applied to obtain final recovered pixels using the refined MV. Experimental results show that the proposed algorithm achieves a significant increase on PSNR and also a great improvement of visual performance. Yang Ling, Yunqiang Liu, Wenjun Zhang 0001 |
ICIP | 4 |
| 2008 | A novel frame recovery algorithm based on spatial and temporal correlationabstractIn video transmission over unreliable channels such as wireless channels or the Internet, transmission errors by for example packet loss are inevitable. For low bit-rate video applications, the loss of a packet may result in the loss of a whole video frame. In this paper, we propose a new algorithm to recover the lost frame. The proposed algorithm explores the spatial and temporal correlation among motion vectors integrative, and then estimates the motion vector field of the lost frame. Finally, adaptive overlapped block motion compensation is used to reconstruct the concealed picture. Results show that the proposed algorithm outperforms other techniques in both PSNR and subjective visual quality. Yang Ling, Yunqiang Liu, Wenjun Zhang 0001 |
ICME | 4 |
| 2008 | Image error-concealment via Block-based Bilateral FilteringabstractWe propose to use Block-based Bilateral Filter (BBF), which extends the classical Bilateral Filter (BF) through operating in block-wise manner, to conceal missing image blocks in the application of compressed image transmission over wireless channels. We show that the problem of error-concealment using BBF can be considered as a superset of image denoising using BF. The BBF has the ability to capture the block-level similarity that well matches the need of error-concealment for block based image compression. Simulation results suggest significant visual and PSNR improvements (up to 11 dB) over various classic and state-of-the-art error-concealment algorithms. Guangtao Zhai, Jianfei Cai 0001, Weisi Lin, Xiaokang Yang 0001, Wenjun Zhang 0001 |
ICME | 5 |
| 2008 | Application of scalable visual sensitivity profile in image and video codingabstractComputational visual attention models have facilitated various aspects of the evolution in visual communication systems. In this paper, we apply the proposed computational model for scalable visual sensitivity profile (SVSP) to image/video processing. SVSP generates a hierarchy of saliency maps that simulate the bottom-up and top-down attention of the human visual system (HVS). It can detect the most significant visual information in a picture and spatially prioritize the information to augment the visual quality, such as in noise-shaping and region-of-interest (ROI) coding of JPEG2000. Furthermore, ROI scalable video coding (SVC) also benefits from the proposed SVSP due to its hierarchical nature, as multiple perceptually prioritized ROIs can be extracted by simply segmenting the SVSP in multiple thresholds in a scalable manner. It can be applied to automatically evaluate the quality for ROI scalability, and thus has the potential to reduce the workload of subjective tests for ROI SVC algorithms in the current MPEG activities. Extensive experiments are conducted to show the performances that justify the effectiveness of the proposed SVSP model. Qian Chen 0024, Guangtao Zhai, Xiaokang Yang 0001, Wenjun Zhang 0001 |
ISCAS | 4 |
| 2008 | Image deringing using quadtree based block-shift filteringabstractIn this paper, we propose an efficient spatial domain deringing algorithm using quadtree (QT) decomposition and block-shift filtering (BSF). The ringing artifacts are located through the QT decomposition of an image down to the 4 × 4 block size. The blocks with suspicious ringing are then replaced with the weighted average of itself and pixel-by-pixel shifted neighboring blocks. The block shifting range is up to 9 × 9 and only those shifted blocks resembling the center one are involved in the averaging stage. Experimental results show that the proposed deringing algorithm can effectively suppress the ringing artifact while preserving image edges and textures well, and the method substantially outperforms most of the existing deringing algorithms reported in the literature, both subjectively and objectively. Guangtao Zhai, Jianfei Cai 0001, Weisi Lin, Xiaokang Yang 0001, Wenjun Zhang 0001 |
ISCAS | 5 |
| 2008 | Cross-dimensional quality assessment for low bitrate videoabstractIn this paper, we address the problem of evaluating perceptual visual quality of compressed video under different settings for mobile communication. A 5-dimensional video feature space is constructed with the codec type, video content, frame size, frame rate and bitrate. The video quality assessment is formulated as a transformation between the multidimensional feature space and the quality space. Subject viewing tests for videos coded with H.263 and H.264 at different bitrates and spatial/temporal resolutions are performed. Based on statistical analysis, we find that the perceptual quality of the decoded video is affected by the encoder type, video content, bitrate, frame rate and frame size, in a descending order for significance. The methodology and findings in this paper designate new researching direction in video quality assessment and can be extended to other applications. Guangtao Zhai, Weisi Lin, Jianfei Cai 0001, Xiaokang Yang 0001, Wenjun Zhang 0001, Minoru Etoh |
ISCAS | 5 |
| 2008 | An Interpolation Based Channel Estimation Method for MIMO OFDM SystemsabstractThis paper proposes a novel channel estimation method to further improve the performance of the optimal pilot sequences in multiple-input multiple-output (MIMO) orthogonal frequency division multiplexing (OFDM) systems. We first assume that the virtual pilot tones superimposed at the data locations over the specific sub-carriers are transmitted from all transmit antennas. We then obtain the virtual received pilot signals at the corresponding locations at the receive antennas by interpolation in the time domain. Finally, the channel parameters are obtained from the combination of the virtual and real received pilot signals over one OFDM symbol based on the least squares (LS) channel estimation. Simulation results show that the proposed channel estimation method provides much better performance than the previous method for the optimal pilot sequences over multiple OFDM symbols, especially in fast time- varying channels. Chengyu Lin 0002, Feng Yang 0006, Wenjun Zhang 0001, Youyun Xu |
VTC Fall | 3 |
| 2008 | Improve throughput of TCP-Vegas in multihop ad hoc networks
Lianghui Ding, Xinbing Wang, Youyun Xu, Wenjun Zhang 0001 |
Comput. Commun. | 4 |
| 2008 | No-reference noticeable blockiness estimation in images
Guangtao Zhai, Wenjun Zhang 0001, Xiaokang Yang 0001, Weisi Lin, Yi Xu 0001 |
Signal Process. Image Commun. | 2 |
| 2008 | Efficient Image Deblocking Based on Postfiltering in Shifted WindowsabstractWe propose a simple yet effective deblocking method for JPEG compressed image through postfiltering in shifted windows (PSW) of image blocks. The MSE is compared between the original image block and the image blocks in shifted windows, so as to decide whether these altered blocks are used in the smoothing procedure. Our research indicates that there exists strong correlation between the optimal mean squared error threshold and the image quality factor Q, which is selected in the encoding end and can be computed from the quantization table embedded in the JPEG file. Also we use the standard deviation of each original block to adjust the threshold locally so as to avoid the over-smoothing of image details. With various image and bit-rate conditions, the processed image exhibits both great visual effect improvement and significant peak signal-to-noise ratio gain with fairly low computational complexity. Extensive experiments and comparison with other deblocking methods are conducted to justify the effectiveness of the proposed PSW method in both objective and subjective measures. Guangtao Zhai, Wenjun Zhang 0001, Xiaokang Yang 0001, Weisi Lin, Yi Xu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2008 | A Cross-Resolution Leaky Prediction Scheme for In-Band Wavelet Video Coding With Spatial ScalabilityabstractIn most existing in-band wavelet video coding schemes, over-complete wavelet transform is used for the motion-compensated temporal filtering (MCTF) of each spatial subband. It can overcome the shift-variance of critical sampling wavelet transform and improve the coding efficiency of the in-band scheme. However, a dilemma exists in the current implementations of in-band MCTF (IBMCTF), which is whether or not to exploit the spatial highpass subbands in motion compensation of the spatial lowpass subband. The absence of the spatial highpass subbands will result in significant quality loss in the reconstructed full-resolution video, whereas the presence of the spatial highpass subbands may bring serious mismatch error in the decoded low-resolution video since the corresponding highpass subbands may be unavailable at the decoder. In this paper, we first analyze the mismatch error propagation in decoding the low-resolution video. Based on our analysis, we then propose a frame-based cross-resolution leaky prediction scheme for IBMCTF. It can make a good tradeoff between alleviating the low-resolution mismatch and improving the full-resolution coding efficiency. Experimental results show that the proposed scheme can dramatically reduce the mismatch error by 0.3-2.5 dB for low resolution, while the performance loss is marginal for high resolution. Wenjun Zhang 0001, Jizheng Xu, Feng Wu 0001, Hongkai Xiong |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2008 | Cross-Dimensional Perceptual Quality Assessment for Low Bit-Rate VideosabstractMost studies in the literature for video quality assessment have been focused on the evaluation of quantized video sequences at fixed and high spatial and temporal resolutions. Only limited work has been reported for assessing video quality under different spatial and temporal resolutions. In this paper, we consider a wider scope of video quality assessment in the sense of considering multiple dimensions. In particular, we address the problem of evaluating perceptual visual quality of low bit-rate videos under different settings and requirements. Extensive subjective view tests for assessing the perceptual quality of low bit-rate videos have been conducted, which cover 150 test scenarios and include five distinctive dimensions: encoder type, video content, bit rate, frame size, and frame rate. Based on the obtained subjective testing results, we perform thorough statistical analysis to study the influence of different dimensions on the perceptual quality and some interesting observations are pointed out. We believe such a study brings new knowledge into the topic of cross-dimensional video quality assessment and it has immediate applications in perceptual video adaptation for scalable video over mobile networks. Guangtao Zhai, Jianfei Cai 0001, Weisi Lin, Xiaokang Yang 0001, Wenjun Zhang 0001, Minoru Etoh |
IEEE Trans. Multim. | 5 |
| 2008 | Efficient Deblocking With Coefficient Regularization, Shape-Adaptive Filtering, and Quantization ConstraintabstractWe propose an effective deblocking scheme with extremely low computational complexity. The algorithm involves three parts: local ac coefficient regularization (ACR) of shifted blocks in the discrete cosine transform (DCT) domain, block-wise shape adaptive filtering (BSAF) in the spatial domain, and quantization constraint (QC) in the DCT domain. The DCT domain ACR suppresses the grid noise (blockiness) in monotone areas. The spatial-domain BSAF alleviates the staircase noise along the edge, and the ringing near the edge and the corner outliers. The narrow quantization constraint set is imposed to prevent possible oversmoothing and improve PSNR performance. Extensive simulation results and comparative studies are provided to justify the effectiveness and efficiency of the proposed deblocking algorithm. Guangtao Zhai, Wenjun Zhang 0001, Xiaokang Yang 0001, Weisi Lin, Yi Xu 0001 |
IEEE Trans. Multim. | 2 |
| 2007 | A Unified Framework for Removing Blocking ArtifactsabstractAl-Fahoum and Reza [1] characterized the blocking artifacts in block based DCT (BDCT) compressed image into five types: grid noise, staircase noise, ringing artifacts, corner outliers and the corruption of edges. Most of the comprehensive deblocking algorithms lack a unified framework, and different artifacts are processed with independent ad hoc schemes. In this paper, we propose a comprehensive postprocessing method for removing all the blocking-related artifacts in the framework of overcomplete wavelet expansion (OWE). We use the wavelet transform modulus maxima extension (WTMME) and angle extracted from the wavelet coefficients of 3-level OWE to represent the image. Both the WTMME and the angle image are reconstructed accordingly using inter-/ intra-band correlation to suppress the influence of the distortions. Simulation and comparative study have demonstrated the effectiveness of the proposed algorithm in terms of both subjective and objective quality of the resultant images. Guangtao Zhai, Wenjun Zhang 0001, Xiaokang Yang 0001, Weisi Lin, Yi Xu 0001 |
ICME | 2 |
| 2007 | Improvement Techniques for the EM-Based Neural Network Approach in RF Components Modeling
Wenjun Zhang 0001, Zhiping Yu |
ISNN (1) | 2 |
| 2007 | A Novel Scheme for Type-II Hybrid ARQ Protocols Using LDPC CodesabstractIn this paper, we propose a novel type-II hybrid ARQ scheme using low density parity check (LDPC) codes. The proposed approach combines low hardware overhead with good decoding performance by using an efficient decoder operating at a much higher rate with a much smaller size parity check matrix. Our scheme makes parts of the previously received data gradually improved until successful decoding of the entire original codeword. We also present a novel efficient framework for constructing rate-compatible LDPC codes applied in our hybrid ARQ scheme. The progressive edge growth (PEG) construction method with zigzag pattern results in linear-time encoding. Xiumin Shi, Shijun Yan, Wenjun Zhang 0001, Yunfeng Guan 0001 |
WCNC | 5 |
| 2007 | An adaptive and fast fractional pixel search algorithm in H.264
Liquan Shen, Zhaoyang Zhang 0002, Zhi Liu 0003, Wenjun Zhang 0001 |
Signal Process. | 4 |
| 2006 | Channel-Aware Frame Dropping for Cellular Video StreamingabstractIn the case of cellular video streaming over wireless channels, burst frame losses may be unavoidable. Considering the unequal importance of different frames in a group-of-pictures (GOP) and the burst-error characteristics of wireless channels, this paper proposes a channel-aware frame dropping scheme so as to shift burst losses into relatively unimportant frames in the same GOP. By using selective retransmission at the radio link layer, a base station can adaptively assign the unequal transmission attempts to different video frames. Simulation results show that the proposed scheme can be aware of the variation of wireless channel conditions, and thus significantly improve error resilience of cellular video streaming Hao Liu 0010, Wenjun Zhang 0001, Songyu Yu, Xiaokang Yang 0001 |
ICASSP (5) | 2 |
| 2006 | Modeling Blocking Visual Sensitivity ProfileabstractBlocking artifact is the most prevailing degradation caused by block-based DCT coding techniques under low bit-rate conditions. To alleviate blockings perceptually, it is desirable to measure the visibility of blocking artifacts. In this paper, we propose an efficient method of estimating the visual sensitivity of blocking artifacts in block-based DCT coding. The differences on block boundaries are measured and transformed into block discontinuity map. We consider the effects of luminance adaptation and texture masking on the blockings and integrate them using nonlinear operator to form an overall masking map. This masking map is then incorporated with the discontinuity map to generate the blocking visual sensitivity map (BVSM). This map can be used to guide perceptual quality assessment, codec parameter optimization, post-processing, etc. We demonstrate the validity of the BVSM through its application in image quality assessment Guangtao Zhai, Wenjun Zhang 0001, Xiaokang Yang 0001, Yi Xu 0001 |
ICME | 2 |
| 2006 | Perceptually-adaptive Motion Compensated Temporal FilteringabstractWe propose a perceptually-adaptive motion compensated temporal filtering (MCTF) method to enhance the visual quality of 3D wavelet video coding schemes with spatial-domain MCTF. In our scheme, a spatio-temporal masking model in image domain is incorporated into the lifting structure of MCTF. The model is used to guide the motion search and the prediction step in MCTF so as to remove the visual redundancy in the video sequence. Experimental results show that the proposed scheme can significantly improve the visual quality of decoded video at different bitrates Wenjun Zhang 0001, Xiaokang Yang 0001, Songyu Yu |
ICME | 3 |
| 2006 | Error-resilience packet scheduling for low bit-rate video streaming over wireless channelsabstractIn the case of low bit-rate video streaming over wireless channels, burst channel errors possibly cause consecutive frame losses. Considering the source-channel characteristics of the specific case, this paper presents a novel error-resilience packet scheduling scheme based on the reference frame selection (RFS) without back-channel, which is called RFS-based packet scheduling (RFS-PS). By applying the delayed packet scheduling scheme in a group-of-picture (GOP) according to different delay requirements, the probability of simultaneous losses between different prediction chains can be reduced. Simulation results show that the proposed RFS-PS scheme can spread out burst packet losses into different chains, and effectively increase error resilience of low bit-rate video streaming over wireless channels Hao Liu 0010, Wenjun Zhang 0001, Xiaokang Yang 0001 |
ISCAS | 2 |
| 2006 | Retransmission-based error spreading for layered video streaming over wireless LANsabstractRobust video streaming over wireless LANs is very challenging due to the burst-error behavior of wireless channels. In this paper, a novel error-resilience scheme, called retransmission-based error spreading, is proposed for layered video streaming over wireless LANs. Considering the unequal importance of different video partitions under the real-time and bandwidth constraints, the proposed scheme uses unequal retransmission attempts at the medium access control (MAC) layer to spread out burst errors. Consequently, burst errors are separated into those relatively unimportant partitions in a picture, and graceful quality degradation can be achieved. Simulation results show that the proposed scheme can be aware of time-varying channel conditions, and thus significantly improve the quality of received video Hao Liu 0010, Wenjun Zhang 0001, Xiaokang Yang 0001 |
ISCAS | 2 |
| 2006 | GES: a new image quality assessment metric based on energy features in Gabor transform domainabstractWe propose Gabor energy similarity (GES), a new full reference image quality assessment metric based on the measuring of Gabor energy features of images. It has been recognized that: 1) 2D Gabor filters can attain the theoretical conjoint resolution limit in spatial and frequency domain defined by Heisenberg's uncertainty principle; 2) it is widely reported that simple cells in visual cortex can be well modeled by 2D Gabor functions; and 3) the feature of Gabor energy is closely related to the model of complex cells in primary visual cortex that are very sensitive to minor changes in nature scene. Motivated by these facts we attempt to design a new image quality assessment by exploring the similarity in Gabor energy functions between the original and the distorted image. The images are firstly decomposed by a filter bank consists of 48 Gabor filters (6 directions, 4 spatial resolutions, and symmetric or anti-symmetric). Then the local energy feature is extracted by computing the modulus of the responses of symmetric and anti-symmetric kernel filters at each point. We finally propose an image quality metric based on averaged cross correlation between two images. Extensive experimental results are used to justify the effectiveness of the proposed GES. Guangtao Zhai, Wenjun Zhang 0001, Xiaokang Yang 0001, Susu Yao, Yi Xu 0001 |
ISCAS | 2 |
| 2006 | Cross-layer conditional retransmission for layered video streaming over cellular networks
Hao Liu 0010, Wenjun Zhang 0001, Xiaokang Yang 0001 |
Comput. Commun. | 2 |
| 2005 | A client-driven scalable cross-layer retransmission scheme for 3G video streamingabstractThe wireless channel is time-varying where burst packet losses often occur during the fading or lossy handovers. In order to avoid unaccepted quality degradation of video streaming over 3G cellular networks, we propose and analyze a client-driven scalable cross-layer (CSC) retransmission scheme. Considering the perceptual importance of different video partitions under the real-time and bandwidth constraints, the proposed scheme uses the radio link-layer retransmission with priority to adapt conventional packet losses in wireless channels; furthermore, it uses the adaptive transport-layer retransmission to provide end-to-end quality-of-service (QoS) guarantees over cellular networks. The simulation experiments show that the proposed scheme can effectively improve the perceptual quality of 3G video streaming as compared to the traditional deadline-based scheme without the prioritized link-layer retransmission. Hao Liu 0010, Wenjun Zhang 0001, Songyu Yu |
ICME | 2 |
| 2005 | Modeling of spiral inductors using artificial neural networkabstractA new model for spiral inductors, which covers wide operation frequency range and full design parameters, is proposed by using artificial neural network (ANN). It is pointed out that a four-layered neural network is superior to a three-layered neural network both on the mapping and generalization abilities in spiral inductor modeling. For the first time, a novel physics-based sampling technique is adopted in modeling procedure. Equipped with this new sampling method, ANN model achieves better speed and accuracy performances though the training data are substantially reduced. The new sampling method can be easily applied to other passive components and be embedded in various modeling frameworks. Wenjun Zhang 0001, Zhiping Yu |
IJCNN | 2 |
| 2005 | Unequal Forced Intra-Refresh for Real-time Multicast VideoabstractIn motion-compensated video coding, the errors caused by packet loss not only impair the reconstruction quality of current frame, but also lead to error propagation to subsequent frames. Based on the error-propagation analysis in a group of pictures (GOP), we propose an unequal forced intra-refresh scheme to increase error resilience of multicast video. According to a GOP-level error-propagation model, the proposed scheme can distribute the unequal number of forced intra-mode MBs to different P-frames of a GOP. Experimental results show that the proposed scheme can effectively mitigate the error-propagation effect and achieve about 0.1~1.1 dB gains over the traditional average scheme in H.264/AVC Hao Liu 0010, Wenjun Zhang 0001, Yutao Dong, Xiangzhong Fang |
MMSP | 2 |
| 2005 | A Concatenated ML Decoder for SFBC-OFDM Systems in Frequency Selective Fading ChannelsabstractThis paper presents a concatenated maximum-likelihood (ML) decoder for space-frequency block coded orthogonal frequency division multiplexing (SFBC-OFDM) systems in severe frequency selective fading channels. The variation of subchannels during a codeword frame will cause the inter-transmit-antenna interference (ITAI), and reduce the diversity gain for SFBC-OFDM systems. The proposed decoder can eliminate the ITAI completely, hence improve the diversity advantage. Simulation results show that the performances of proposed decoder is very close to that of the optimal ML decoder in highly delay spread channels. However, the complexity of proposed decoder is much lower than that of the optimal ML decoder Xiaodong Zhang 0001, Wenjun Zhang 0001 |
PIMRC | 4 |
| 2005 | Ordered Statistics Based Rate Allocation Scheme for Closed Loop MIMO-OFDM SystemabstractWe propose a new rate allocation algorithm for closed loop MIMO-OFDM system. The new scheme utilizes ordered statistics of channel matrix's singular value and is significantly simpler than ideal scheme which uses water-filling in both frequency and space domain. The proposed scheme is more flexible in that it does not force the use of large antenna array as "conventional frequency flat constraint" scheme does [J.R.B. Joon Hyun Sung, January 2003]. It also outperforms FFC algorithm under equal hardware complexity. Siya Wang, Wenjun Zhang 0001, Xiaoying Gan |
WiOpt | 3 |
| 2004 | Solving Engineering Design Problems by Social Cognitive Optimization
Xiao-Feng Xie 0001, Wenjun Zhang 0001 |
GECCO (1) | 2 |
| 2001 | Realization of semiconductor device synthesis with the parallel genetic algorithmabstractIn this presented paper, to accomplish semiconductor device synthesis for TCAD application, the Parallel Genetic Algorithm is applied as the core searching algorithm for "acceptability region" of device designables, which satisfy the designed device performance. The results of some experiments on FIBMOS are shown, which indicate the Parallel Genetic Algorithm is an efficient and fast searching algorithm to fulfill device synthesis. Some potential problems related to device synthesis are also discussed. Xiao-Feng Xie 0001, Wenjun Zhang 0001, Zhilian Yang |
ASP-DAC | 3 |
| 1999 | Enhancing the Efficiency of Reduction of Large RC networks By Pole Analysis via Congruence TransformationsabstractAmong the RC reduction algorithms, the algorithm of PACT (Pole Analysis via Congruence Transformations) has been proved to have several advantages. However, the original implementation of the algorithm destroys the sparsity of the internal capacitance matrix. Consequently, the LASO process, used in the computation of the dominant eigenvalues and eigenvectors, becomes very time-consuming. Therefore, the efficiency of the algorithm needs to be improved. In this paper, a new method to implement the PACT algorithm is presented. In order to maintain the sparsity of the matrices, we use a special Lanczos algorithm to directly compute the eigenvalues and eigenvectors by solving a large sparse symmetric generalized eigenvalue problem, At the same time, this approach can avoid some matrix multiplication to speed up the reduction process. We have constructed a RC reduction tool with the new implementation method. The application of the tools to several RC networks has shown that this tool greatly outperforms the original implementation. Wenjun Zhang 0001, Lilin Tian, Zhilian Yang |
ASP-DAC | 2 |