EDBT 2026 Demo / reviewers in the wild / expert
Suhua Tang
dblp:70/3828
· DBLP profile ↗
65ranked-venue papers
23as first author
22since 2021 · last 2026
0000-0002-5784-8411ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 26 · 14 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 4 since 2021Artificial intelligence and machine learning · 8 · 2 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Inferring Targets from Calibrated Hesitations via Mutual Information Maximization in Multi-Behavior RecommendationabstractMulti-behavior recommendation enriches user preference modeling by incorporating diverse auxiliary interactions. However, most existing methods simply treat interactions without target behaviors as absolute negative feedback. This strategy ignores an important intermediate state known as user hesitation, where users exhibit strong intent but fail to complete the final conversion due to various reasons. Consequently, models cannot distinguish true disinterest from intended but hesitant behavior, which introduces substantial noise into preference modeling. To address this issue, we propose a novel framework named Calibrated Hesitation Analysis for Multi-Behavior Recommendation via Mutual Information Maximization (CHARM). Specifically, we aggregate auxiliary behaviors that lead to successful conversions into latent intent representations and train an inference network by maximizing the mutual information between these intents and observed target behaviors. We then apply this network to auxiliary behaviors without conversion, under the assumption that the conversion had occurred, in order to infer latent conversion probabilities and identify high-intent hesitation candidates. Furthermore, to distinguish genuine hesitation from interaction termination caused by competing item choices, we design a competitor substitution penalty strategy to refine hesitation confidence scores. Finally, the calibrated hesitation set is incorporated into the recommendation process to improve ranking quality. Extensive experiments on three real-world datasets demonstrate that CHARM consistently outperforms existing state-of-the-art methods. The source code is available at https://github.com/city59/CHARM. Cheng Li 0058, Yong Xu 0001, Suhua Tang, Xin He 0017, Jinde Cao |
SIGIR | 3 |
| 2026 | Long-Tail Inductive Feature Transfer for Knowledge-Aware Multi-Behavior RecommendationabstractMulti-Behavior recommendation captures fine-grained user intents by jointly modeling multiple interaction behaviors, significantly improving recommendation accuracy and diversity. However, existing methods often overlook complex semantic relations between items and entities. Moreover, constructing separate subgraphs for different interaction types frequently results in sparse graph structures, which exacerbates the long-tail problem. To address these challenges, we propose Long-tail Inductive Feature Transfer (LIFT), a method for knowledge-aware multi-behavior recommendation. To mitigate uneven node distribution in graph structures, we introduce a knowledge transfer-based feature reconstruction mechanism. Specifically, we first drop a portion of the neighbors of head nodes to construct proxy representations for tail nodes, training a reconstructor on the tail proxy to reconstruct the original head node features. The trained reconstructor is then used to backfill missing neighbor information for tail nodes, thereby achieving a more balanced feature distribution across nodes. Furthermore, we integrate multi-behavior and semantic contrastive learning to jointly optimize the representations. Extensive experiments on four datasets demonstrate that LIFT outperforms state-of-the-art methods, with further analysis validating its uniformity in representation learning. The source code is available at:https://github.com/city59/LIFT. Cheng Li 0058, Yong Xu 0001, Suhua Tang, Weiguo Wang, Xin He 0017, Jinde Cao |
IEEE Trans. Big Data | 3 |
| 2026 | Coupled Phase-Shift STAR-RIS Enabled Integrated Over-the-Air Computation and CommunicationsabstractTo meet the emerging demands for rapid data aggregation and reliable information transmission in future wireless applications, a novel system architecture integrating Over-the-Air Computation (AirComp) and downlink multi-user communication via a simultaneously transmitting and reflecting reconfigurable intelligent surface (STAR-RIS) is proposed in this paper. In the considered cellular scenario, an unmanned aerial vehicle (UAV) carries a STAR-RIS beneath its fuselage, creating a programmable aerial platform that concurrently serves Internet-of-Things (IoT) devices and conventional mobile users. The STAR-RIS operates in transmission mode to enable efficient wireless data aggregation of IoT devices, while its reflection mode establishes high-quality downlink channels from the base station (BS) to multiple users. Capturing the true electromagnetic behavior of the STAR-RIS, we explicitly model the practical coupling between the reflection and transmission phase shifts. Two optimization problems are then formulated: one minimizes AirComp distortion and the other maximizes the minimum user rate in the downlink. Both non-convex problems are tackled by efficient iterative algorithms derived from the penalty dual decomposition (PDD) framework. Extensive simulations confirm that the proposed design markedly outperforms baseline approaches and its performance can approach that of ideal phase-shift control by enhancing the key system parameters. Additionally, the trade-off between computation and communication performance is demonstrated. Shuzhen Yuan, Chao Zhang 0003, Junjie Fang, Yuanwei Liu, Suhua Tang, Qingqing Wu 0001 |
IEEE Trans. Commun. | 5 |
| 2025 | Investigation of the impact of Doppler shift on OFDM-based rangingabstractThe performance of GNSS based positioning, not only the accuracy but also the outage (where there is no sufficient number of satellites for position computation), is greatly degraded in urban canyons. To solve this problem, previous works have exploited vehicles and road side units as anchors, and suggested estimating the distance using phase information of V2X (vehicle-to-everything) OFDM signals. To further reduce the outage and improve the positioning accuracy, it is reasonable to use as positioning anchors the large number of LEO (Low Earth Orbit) satellites, which currently are mainly deployed for ubiquitous communications. Experiments in some literature revealed that LEO satellites use an OFDM-like structure to transmit data, and it is possible to exploit the OFDM phase information of LEO satellite signals to estimate the distance. It is well known that the impact of Doppler shift is serious at high speeds. To verify the feasibility of using the OFDM-based ranging with LEO satellites, this paper analyzes the impact of Doppler shift on the phase measurement of OFDM signals, and reports the ranging performance under 3D ray tracing simulations. Suhua Tang, Sadao Obana |
VTC2025-Fall | 1 |
| 2025 | Enhancing semantic audio-visual representation learning with supervised multi-scale attention
Jiwei Zhang 0012, Yi Yu 0001, Suhua Tang, Guo-Jun Qi, Haiyuan Wu, Hirotaka Hachiya |
Pattern Anal. Appl. | 3 |
| 2025 | WiCG: Heartbeat Sensing Using COTS WiFi Devices with Common AntennaabstractVital sign detection, based on Channel State Information (CSI) from commercial off-the-shelf (COTS) WiFi devices, has become a popular research area. Previous works in this field mainly focus on respiration, while heartbeat sensing has not been well studied yet, because its signal is very weak and overwhelmed by hardware noises and the respiration signal. Different from existing research that exploits directional antenna, the proposed WiCG ( Wi Fi C ardio G ram) system uses common antennas, and not only accurately senses heartbeat rate but also provides the heartbeat signal for further analysis in complex real-life home scenes. Specifically, we first propose an effective denoising solution for Wi-Fi CSI by exploiting its spatial structure, which exhibits strong correlation among the In-phase/Quadrature components. Leveraging this characteristic with Principal Component Analysis (PCA) achieves effective reduction of ambient noise in both the amplitude and phase of the CSI. Then, we introduce a heartbeat enhancement scheme that utilizes the periodicity of the heartbeat signal. By applying Singular Spectrum Analysis (SSA), the complex effects of residual noise and respiratory interference are effectively mitigated. Extensive experiments have proven that WiCG can effectively sense the heartbeat rate. In a real deployment environment, the average detection error can be reduced to 0.28 bpm, close to current commercial heartbeat sensors. Zhi Liu 0002, Celimuge Wu, Jie Li 0002, Suhua Tang |
ACM Trans. Sens. Networks | 5 |
| 2024 | Enhancing the Reliability of NR-V2X Sidelink Broadcast through RelayabstractNR-V2X sidelink (SL) broadcast is used for the real-time exchange of position and other information between adjacent vehicles. However, its reliability degrades much in non-line-of-sight environments, where Blind ReTransmission (BReTX) does not work well. Relay can effectively address this issue, and a relaybased SL broadcast method was proposed in previous work. However, in this method, each vehicle must detect communication Link Quality (LQ) and exchange LQ Indicators (LQIs) with its neighbors to select a proper relay, which causes much overhead. This paper aims to improve the relay-based SL broadcast from two aspects. (i) To reduce the overhead, we suggest aggregating LQIs per direction, making the overhead of sharing LQIs irrelevant to the number of vehicles. (ii) To further improve reliability, we use an explicit ACK to deal with the potential failure of the transmission from the source vehicle to the relay. Using explicit ACKs allows the source vehicle to retransmit the packet until the relay vehicle correctly receives it, after which the relay vehicle handles the remaining retransmissions. Simulation results confirm that the proposed method, RReTX-ACK, improves the packet dissemination rate by approximately 8.30% within a 200 m distance from the source vehicle in an intersection scenario, compared to BReTX without using relay vehicles. Shiro Aoki, Suhua Tang |
APCC | 2 |
| 2024 | Combining AF-based Relay and Multiantenna for Over-the-Air ComputationabstractOver-the-air computation has attracted much attention in wireless data aggregation, which collects and processes data simultaneously with high efficiency. Different propagation distances from nodes to the sink or channel fading may lead to misalignment in signal magnitude and result in computation error. In our previous work, amplify and forward (AF) based relay has been investigated. With multiantenna available at the sink, it is possible to achieve better performance by the misalignment allowed optimization (Miso), compared with zero forcing (ZF) which requires all signals aligned in signal magnitude. But Miso does cause many signals misaligned in signal magnitude, leading to a biased result. Aiming to further improve system performance, in this work, we study how to combine the AF based relay with Miso by a joint optimization, and solve it with an iterative algorithm. Extensive simulations verify the effectiveness of the proposed method. Suhua Tang, Sadao Obana |
IWCMC | 1 |
| 2024 | Evaluation of Distance Estimation Using Phase Information of OFDM SignalabstractIn urban areas, due to the obstruction/reflection of GNSS signals by roadside buildings, position accuracy is greatly degraded, or even outage occurs when no enough satellites are available. To address this problem, vehicles and roadside units have been used as positioning anchors to estimate pedestrian position. Although previous methods exploited channel state information, which is superior to received signal strength for distance estimation, they rely on signal attenuation characteristics, and the accuracy is limited by temporal resolution and multipath propagation. This work focuses on using the phase information of V2X OFDM signals to estimate distance differences more robustly. We investigated the impact of the sampling rate, a crucial parameter, and conducted simulation evaluations using 3D maps and ray-tracing. Additionally, we built a testbed and performed evaluations with real-world OFDM signal data, which demonstrates that phase information can significantly improve the accuracy of distance measurements. Noriyasu Kikuchi, Nobuaki Kubo, Suhua Tang |
VTC Fall | 4 |
| 2024 | Miso: Misalignment Allowed Optimization for Multiantenna Over-the-Air ComputationabstractOver-the-air computation (AirComp), as an effective method to wireless data aggregation, has attracted much attention recently. It helps to improve network efficiency and scalability by integrating communication and computation in the air. In AirComp, both signal magnitude misalignment and noise lead to computation error. In the single antenna case, by allowing misalignment in signal magnitude, a good tradeoff can be achieved between signal distortion and noise power, which leads to a minimal error. In the multi-antenna case, usually the zero-forcing policy is used to enforce signal magnitude alignment (no distortion), which, however, increases noise and affects the overall computation error. To better exploit multiple antennas at the sink, in this paper, we propose a misalignment allowed optimization (Miso) method for AirComp. Specifically, a group of nodes whose signals may be misaligned are dynamically selected, and other signals are aligned to a higher level with higher quality. On this basis, the optimization of multi-antenna AirComp is converted to a difference of convex problem and is solved iteratively. Simulations confirm that the proposed method greatly reduces computation error and scales better with the number of nodes, compared with previous methods. Suhua Tang, Chao Zhang 0003, Jie Li 0002, Sadao Obana |
IEEE Internet Things J. | 1 |
| 2023 | AoA Estimation for High Accuracy BLE PositioningabstractPositioning information plays an important role in various indoor services such as route navigation and robot control. Bluetooth Low Energy (BLE), with low power consumption and low cost, are widely installed in mobile devices, and used for indoor localization. Previous methods usually exploit received signal strength indicator (RSSI) to compute position, and the performance is limited. To further improve positioning accuracy, it is desirable to use angle-of-arrival (AoA) information together. However, a BLE receiver has only one receiving module, and cannot receive signals from multiple antennas simultaneously. To solve this problem, Bluetooth 5.1 added the Constant Tone Extension (CTE) to the end of a BLE packet. Then, it is possible for a BLE receiver to switch the antennas and acquire phase information per antenna successively. In this paper, by adjusting parameters related to CTE, we first investigate how to compute AoA by linear regression. Then, we further discuss how to estimate AoA by exploiting the Multiple Single Classification (MUSIC) method. Experimental results show that linear regression can achieve an AoA error of 5.7 degrees, while using MUSIC helps to reduce this error to 2.3 degrees, which is promising for high accuracy positioning. Yuto Yamami, Suhua Tang |
CCNC | 2 |
| 2023 | Reliable NR-V2X Broadcast Transmission by RelayabstractVehicular communication plays an important role in cooperative perception. 3GPP has successively standardized LTE-V2X and NR-V2X to support different kinds of vehicular communications, especially local information dissemination among vehicles by broadcast. There is no feedback for broadcast transmission. Therefore, blind retransmission usually is used to improve system reliability. But its effect is limited in the non-line-of-sight (NLOS) scenario. In this paper, with NR-V2X mode 2 as an example, we study how to improve the reliability of local broadcast transmission by using relay, from two aspects. (i) A vehicle allocates resources for the retransmissions at a relay vehicle in advance to reduce the delay. (ii) A relay vehicle is selected to help multiple vehicles simultaneously, considering the link quality between vehicles. Simulation evaluations confirm that using relay helps to improve information dissemination at the NLOS intersections, even when there is no vehicle at the intersection center. Suhua Tang, Sadao Obana |
VTC Fall | 1 |
| 2023 | Multi-scale network with shared cross-attention for audio-visual correlation learning
Jiwei Zhang 0012, Yi Yu 0001, Suhua Tang, Wei Li 0012 |
Neural Comput. Appl. | 3 |
| 2023 | Variational Autoencoder with CCA for Audio-Visual Cross-modal RetrievalabstractCross-modal retrieval is to utilize one modality as a query to retrieve data from another modality, which has become a popular topic in information retrieval, machine learning, and databases. Finding a method to effectively measure the similarity between different modality data is the major challenge of cross-modal retrieval. Although several research works have calculated the correlation between different modality data via learning a common subspace representation, the encoder’s ability to extract features from multi-modal information is not satisfactory. In this article, we present a novel variational autoencoder architecture for audio–visual cross-modal retrieval by learning paired audio–visual correlation embedding and category correlation embedding as constraints to reinforce the mutuality of audio–visual information. On the one hand, audio encoder and visual encoder separately encode audio data and visual data into two different latent spaces. Further, two mutual latent spaces are respectively constructed by canonical correlation analysis. On the other hand, probabilistic modeling methods are used to deal with possible noise and missing information in the data. Additionally, in this way, the cross-modal discrepancies from intra-modal and inter-modal information are simultaneously eliminated in the joint embedding subspace. We conduct extensive experiments over two benchmark datasets. The experimental results confirm that the proposed architecture is effective in learning audio–visual correlation and is appreciably better than the existing cross-modal retrieval methods. Jiwei Zhang 0012, Yi Yu 0001, Suhua Tang, Wei Li 0012 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2023 | Melody Generation from Lyrics with Local InterpretabilityabstractMelody generation aims to learn the distribution of real melodies to generate new melodies conditioned on lyrics, which has been a very interesting topic in the area of artificial intelligence and music. However, a challenging issue still limits the quality and reliability of melody generation conditioned on lyrics: how to enhance the interpretability between the input lyrics and generated melodies so humans can understand their relationships. To solve this issue, in this article, we propose a model for melody generation from lyrics with local interpretability, which contains two significant contributions: (i) Mutual information between input lyrics and generated melody is exploited to instruct the training of the network, which avoids the loss of content consistency during the training stage. (ii) Transformer is explored to efficiently extract semantic features from lyrics sequences, which provides more interpretable correlations between different syllables in lyrics. Experiments on a large-scale dataset with paired lyrics-melodies demonstrate that the proposed approach can generate higher-quality melodies from lyrics compared with existing methods. Wei Duan 0004, Yi Yu 0001, Xulong Zhang 0001, Suhua Tang, Wei Li 0012, Keizo Oyama |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | Multi-Slot Over-the-Air Computation in Fading ChannelsabstractIoT systems typically involve separate data collection and processing, and face the scalability issue when the number of nodes increases. For some tasks, only the result of data fusion is needed. Then, the whole process can be realized in an efficient way, integrating the data collection and fusion in one step by over-the-air computation (AirComp). Its shortcoming, however, is signal distortion when channel gains of nodes are different, which cannot be well solved by transmission power control alone in times of deep fading. To address this issue, in this paper, we propose a multi-slot over-the-air computation (MS-AirComp) framework for the sum estimation in fading channels. Compared with conventional data collection (one slot for each node) and AirComp (one slot for all nodes), MS-AirComp is an alternative policy that lies between them, exploiting multiple slots to improve channel gains so as to facilitate power control. Avoiding to obtain instantaneous channel gains of all nodes at the sink is a key point. Specifically, the transmissions are distributed over multiple slots and a threshold of channel gain is set for distributed transmission scheduling. Each node transmits its signal only once, in the slot when its channel gain first gets above the threshold, or in the last slot when its channel gain remains below the threshold. Theoretical analysis gives the closed-form of the computation error in fading channels, based on which the optimal parameters are found. Noticing that computation error tends to be reduced at the cost of more transmission power, a method is suggested to control the increase of transmission power. Simulations confirm that the proposed method can effectively reduce computation error, compared with state-of-the-art methods. Suhua Tang, Petar Popovski, Chao Zhang 0003, Sadao Obana |
IEEE Trans. Wirel. Commun. | 1 |
| 2022 | LSTM-Based High Precision Pedestrian PositioningabstractPrevention of pedestrians traffic accidents has become an important issue in intelligent transportation system. Pedestrian to vehicle communication, in which pedestrians send their position information to surrounding vehicles, is effective in reducing accidents when pedestrians are out of sight of vehicles. However, in urban canyons, the precision of pedestrian positioning via GPS may be significantly degraded due to obstructions and reflections of roadside buildings. To solve this problem, our previous work has suggested using vehicles as anchors for pedestrian positioning, and estimating pedestrian-vehicle distance from instantaneous channel state information (CSI), by using Support Vector Regression (SVR). In this paper, based on the fact that pedestrian-vehicle distance changes continuously, we propose to estimate the distance from a CSI sequence by using an LSTM (Long short-term memory) model, and solve the problem of packet loss that may affect the CSI collection. Simulations using 3D ray tracing show that the proposed method reduces the distance error and the horizontal positioning error by 24.8 % and 41.7%, respectively, compared to the previous method using SVR. Masaki Inoue, Suhua Tang, Sadao Obana |
CCNC | 2 |
| 2022 | Deep Attention-Based Alignment Network for Melody Generation from Incomplete LyricsabstractWe propose a deep attention-based alignment network, which aims to automatically predict lyrics and melody with given incomplete lyrics as input in a way similar to the music creation of humans. Most importantly, a deep neural lyrics-to-melody net is trained in an encoder-decoder way to predict possible pairs of lyrics-melody when given incomplete lyrics (few keywords). The attention mechanism is exploited to align the predicted lyrics with the melody during the lyrics-to-melody generation. The qualitative and quantitative evaluation metrics reveal that the proposed method is indeed capable of generating proper lyrics and corresponding melody for composing new songs given a piece of incomplete seed lyrics. Gurunath Reddy M, Zhe Zhang 0050, Yi Yu 0001, Florian Harscoët, Simon Canales, Suhua Tang |
ISM | 6 |
| 2022 | Melody Generation from Lyrics Using Three Branch Conditional LSTM-GAN
Wei Duan 0004, Rajiv Ratn Shah, Suhua Tang, Wei Li 0012, Yi Yu 0001 |
MMM (1) | 5 |
| 2022 | Exploiting Phase Difference of Arrival of V2X Signals for Pedestrian PositioningabstractPosition information plays an important role in preventing pedestrian accidents by pedestrian-to-vehicle communication. But the computation of pedestrian position, based on GPS, may fail in urban canyons due to the obstruction of roadside building. This problem can be solved by using vehicles and roadside units as anchors for pedestrian positioning, where trilateration is used to compute pedestrian position based on distances to anchors. But the performance is degraded by multipath propagation and limited by the time resolution. To address this problem, in this paper, we investigate how phase information of OFDM signals in V2X communications varies with the propagation distance, and exploit the phase difference of arrival to estimate the distance difference. We study how to deal with the inter-symbol interference in OFDM symbols and combine multiple estimations of distance difference to improve the accuracy. Simulation evaluations by 3D ray-tracing confirm the effectiveness of the proposed method. Suhua Tang, Sadao Obana |
VTC Fall | 1 |
| 2022 | Unified Performance Analysis of Stochastic Clustered Cooperative Systems With Distance-Based Relay SelectionabstractIn this paper, we present a unified performance analysis of stochastic clustered cooperative systems with distance-based relay selection, where the destination is located at the cluster center, the locations of the source and the candidate relays follow an independent and identical distribution, and there exists a Poisson field of interferers. As the distance-based relay selection leads to spatial correlation between cooperative transmission distances in a cluster, we first derive a general joint probability distribution function of cooperative transmission distances based on$\psi $-order statistics. Then, the transmission success probability of the stochastic clustered cooperative systems with decode-and-forward (DF) scheme is derived. In order to reduce computational complexity, we offer an approximate expression of the transmission success probability. Furthermore, we provide an asymptotic transmission success probability in the interference-limited scenario, when the number of candidate relays in a cluster is sufficiently large. Besides, the performance analysis framework is extended into the clustered cooperative systems with finite relay distribution regions. Taking the cluster member distributions of Thomas cluster process (TCP) and Matérn cluster process (MCP) as examples, we verify our theoretical analysis via Monte-Carlo simulations. With the theoretical results, the optimization of relay selection parameter is also demonstrated. Fangzhou Yu, Chao Zhang 0003, Suhua Tang |
IEEE Trans. Wirel. Commun. | 3 |
| 2021 | IMP: Impedance Matching Enhanced Power-Delivered-to-Load Optimization for Magnetic MIMO Wireless Power Transfer SystemabstractRecently, multiple-input multiple-output (MIMO) technology has been introduced into magnetic resonant coupling (MRC) enabled wireless power transfer (WPT) systems for concurrent charging of multiple devices. However, impedance mismatching phenomena caused by strong TX-RX or RX-RX coupling greatly affect the power delivered to load (PDL) in practical charging systems. To solve this issue, we propose an effective scheduling algorithm for Impedance Matching enhanced PDL optimization in MIMO MRC-WPT systems (called IMP), which integrates the transmitter scheduling together with the impedance matching techniques, i.e., adjusting TX coils for tuning TX-RX coupling and grouping RXs to separate strongly coupled RX pairs. We formulate this as a joint optimization problem and decouple it into three sub-problems, i.e., current scheduling, coil adjustment, and RX grouping, and solve them through alternating direction method of multipliers (ADMM) based, tabu search (TS) based, and graph clique cover based algorithms, respectively. Extensive experiments are performed on a prototype testbed, and the results demonstrate the effectiveness of our solution. Compared with the state-of-the-art power transfer efficiency (PTE) maximization solution, the proposed algorithm IMP achieves a 74.7X performance improvement of PDL on average. Wangqiu Zhou, Hao Zhou 0001, Wenxiong Hua, Fengyu Zhou 0003, Xiang Cui, Suhua Tang, Zhi Liu 0002, Xiang-Yang Li 0001 |
IWQoS | 6 |
| 2020 | HPNet: A Compressed Neural Network for Robust Hybrid Precoding in Multi-User Massive MIMO SystemsabstractIn multi-user millimeter wave (mmWave) communications, massive multiple-input multiple-output (MIMO) systems can achieve high gain and spectral efficiency significantly. To reduce the hardware complexity and energy consumption of massive MIMO systems, hybrid precoding as a crucial technique has attracted extensive attention. Most previous works for hybrid precoding developed algorithms based on optimization or exhaustive search approaches that either lead to sub-optimal performance or have high computational complexity. Motivated by the thought of cross-fertilization between Data-driven and Model-driven approaches, we consider applying deep learning approach and introduce the Hybrid Precoding Network(HPNet), which is a compressed deep neural network exploiting the feature extracting (thanks to convolutional kernels) and generalization ability of neural networks and the natural sparsity of mmWave channels. The HPNet takes imperfect channel state information (CSI) as the input and predicts the analog precoder and baseband precoder for multi-user massive MIMO systems. Moreover, in order to make the approach more practical in real scenarios, we further introduce a model compression algorithm, using network pruning, to greatly reduce the computational complexity and memory usage of the neural network while almost retaining the model performance and then assess the influence of pruned parameters in the network. Numerical experiments demonstrate that HPNet outperforms state-of-the-art hybrid precoding schemes with higher performance and stronger robustness. Finally, we analyze and compare the computational complexity of different schemes. Mingyang Chai, Suhua Tang, Ming Zhao 0008, Wuyang Zhou |
GLOBECOM | 2 |
| 2020 | Lyrics-Conditioned Neural Melody Generation
Yi Yu 0001, Florian Harscoët, Simon Canales, Gurunath Reddy M, Suhua Tang, Junjun Jiang |
MMM (2) | 5 |
| 2020 | Dynamic Control of Transmission Interval for Efficient Pedestrian-to-Vehicle Communication Based on Channel Utilization RateabstractPedestrian-to-vehicle (P2V) communication, which sends information from pedestrians to nearby vehicles by wireless signals, has attracted much attention recently. In P2V communication, when many pedestrians share the same channel and transmit frequently, congestion occurs, which not only increases the delay but also degrades packet delivery rate. In our previous work, a method was proposed to estimate the degree of risk of pedestrians according to the context information of pedestrians (e.g., pedestrian property, distance to vehicles, etc.), and on this basis differentiate pedestrians by assigning different transmission priorities to pedestrians, with a lower priority leading to a longer transmission interval. However, this method does not consider actual channel utilization and cannot well adapt to different scenarios. In this paper, we extend our previous method, and use channel utilization rate as the feedback of the transmission control. This feedback is essential in making the algorithm track the changes in the number of pedestrians and their degrees of risk. Evaluations on network simulator confirm that the proposed method effectively improves packet delivery rate by 12.3% on average meanwhile retaining the same delay compared with the previous method. Shun Ito, Suhua Tang, Sadao Obana |
VTC Spring | 2 |
| 2020 | Exploiting Large Vehicles with High Antenna for Efficient Relay in Inter-Vehicle CommunicationabstractInter-vehicle communication (IVC) helps to prevent vehicle accidents by disseminating vehicle position and speed information, and road/traffic information among vehicles. Relay vehicle selection control (for disseminating road/traffic information in a wide range) and prioritized transmission control (for disseminating fast and reliably urgent information such as sudden brake information) play important roles in IVC. Our previous work proposed to integrate the two control functions explicitly in the MAC (Media Access Control) layer. Relay vehicle selection control is realized by sorting vehicles in the descending order of their distances from the transmitting vehicle and on this basis a different waiting time is set to each vehicle via using CW (Contention Window) in the EDCA (Enhanced Distributed Channel Access) mechanism of wireless LANs. In this paper, we further consider the coexistence of ordinary and large vehicles. A large vehicle, with a high antenna and a long transmission distance, is more preferentially selected as relay by introducing a new metric-potential relay distance. Simulation results show that by preferentially selecting large vehicles as relay, the proposed method can improve the reachability by up to 10% and reduce the delay by 35% in the highway scenario, and can improve the reachability by up to 4% and reduce the delay by up to 13% in the urban scenario. Takuya Mori, Suhua Tang, Sadao Obana |
VTC Spring | 2 |
| 2020 | Context-Patch Face Hallucination Based on Thresholding Locality-Constrained Representation and Reproducing LearningabstractFace hallucination is a technique that reconstructs high-resolution (HR) faces from low-resolution (LR) faces, by using the prior knowledge learned from HR/LR face pairs. Most state-of-the-arts leverage position-patch prior knowledge of the human face to estimate the optimal representation coefficients for each image patch. However, they focus only the position information and usually ignore the context information of the image patch. In addition, when they are confronted with misalignment or the small sample size (SSS) problem, the hallucination performance is very poor. To this end, this paper incorporates the contextual information of the image patch and proposes a powerful and efficient context-patch-based face hallucination approach, namely, thresholding locality-constrained representation and reproducing learning (TLcR-RL). Under the context-patch-based framework, we advance a thresholding-based representation method to enhance the reconstruction accuracy and reduce the computational complexity. To further improve the performance of the proposed algorithm, we propose a promotion strategy called reproducing learning. By adding the estimated HR face to the training set, which can simulate the case that the HR version of the input LR face is present in the training set, it thus iteratively enhances the final hallucination result. Experiments demonstrate that the proposed TLcR-RL method achieves a substantial increase in the hallucinated results, both subjectively and objectively. In addition, the proposed framework is more robust to face misalignment and the SSS problem, and its hallucinated HR face is still very good when the LR test face is from the real world. The MATLAB source code is available at https://github.com/junjun-jiang/TLcR-RL. Junjun Jiang, Yi Yu 0001, Suhua Tang, Jiayi Ma 0001, Akiko Aizawa, Kiyoharu Aizawa |
IEEE Trans. Cybern. | 3 |
| 2020 | Ensemble Super-Resolution With a Reference DatasetabstractBy developing sophisticated image priors or designing deep(er) architectures, a variety of image super-resolution (SR) approaches have been proposed recently and achieved very promising performance. A natural question that arises is whether these methods can be reformulated into a unifying framework and whether this framework assists in SR reconstruction? In this paper, we present a simple but effective single image SR method based on ensemble learning, which can produce a better performance than that could be obtained from any of SR methods to be ensembled (or called component super-resolvers). Based on the assumption that better component super-resolver should have larger ensemble weight when performing SR reconstruction, we present a maximum a posteriori (MAP) estimation framework for the inference of optimal ensemble weights. Especially, we introduce a reference dataset, which is composed of high-resolution (HR) and low-resolution (LR) image pairs, to measure the SR abilities (prior knowledge) of different component super-resolvers. To obtain the optimal ensemble weights, we propose to incorporate the reconstruction constraint, which states that the degenerated HR estimation should be equal to the LR observation one, as well as the prior knowledge of ensemble weights into the MAP estimation framework. Moreover, the proposed optimization problem can be solved by an analytical solution. We study the performance of the proposed method by comparing with different competitive approaches, including four state-of-the-art nondeep learning-based methods, four latest deep learning-based methods, and one ensemble learning-based method, and prove its effectiveness and superiority on some general image datasets and face image datasets. Junjun Jiang, Yi Yu 0001, Zheng Wang 0007, Suhua Tang, Ruimin Hu, Jiayi Ma 0001 |
IEEE Trans. Cybern. | 4 |
| 2019 | Face hallucination through differential evolution parameter map learning with facial structure prior
Junjun Jiang, Jiayi Ma 0001, Suhua Tang, Yi Yu 0001, Kiyoharu Aizawa |
Inf. Sci. | 3 |
| 2019 | Category-Based Deep CCA for Fine-Grained Venue Discovery From Multimodal DataabstractIn this work, travel destinations and business locations are taken as venues. Discovering a venue by a photograph is very important for visual context-aware applications. Unfortunately, few efforts paid attention to complicated real images such as venue photographs generated by users. Our goal is fine-grained venue discovery from heterogeneous social multimodal data. To this end, we propose a novel deep learning model, category-based deep canonical correlation analysis. Given a photograph as input, this model performs: 1) exact venue search (find the venue where the photograph was taken) and 2) group venue search (find relevant venues that have the same category as the photograph), by the cross-modal correlation between the input photograph and textual description of venues. In this model, data in different modalities are projected to a same space via deep networks. Pairwise correlation (between different modality data from the same venue) for exact venue search and category-based correlation (between different modality data from different venues with the same category) for group venue search are jointly optimized. Because a photograph cannot fully reflect rich text description of a venue, the number of photographs per venue in the training phase is increased to capture more aspects of a venue. We build a new venue-aware multimodal data set by integrating Wikipedia featured articles and Foursquare venue photographs. Experimental results on this data set confirm the feasibility of the proposed method. Moreover, the evaluation over another publicly available data set confirms that the proposed method outperforms state of the arts for cross-modal retrieval between image and text. Yi Yu 0001, Suhua Tang, Kiyoharu Aizawa, Akiko Aizawa |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2019 | Deep Cross-Modal Correlation Learning for Audio and Lyrics in Music RetrievalabstractDeep cross-modal learning has successfully demonstrated excellent performance in cross-modal multimedia retrieval, with the aim of learning joint representations between different data modalities. Unfortunately, little research focuses on cross-modal correlation learning where temporal structures of different data modalities, such as audio and lyrics, should be taken into account. Stemming from the characteristic of temporal structures of music in nature, we are motivated to learn the deep sequential correlation between audio and lyrics. In this work, we propose a deep cross-modal correlation learning architecture involving two-branch deep neural networks for audio modality and text modality (lyrics). Data in different modalities are converted to the same canonical space where intermodal canonical correlation analysis is utilized as an objective function to calculate the similarity of temporal structures. This is the first study that uses deep architectures for learning the temporal correlation between audio and lyrics. A pretrained Doc2Vec model followed by fully connected layers is used to represent lyrics. Two significant contributions are made in the audio branch, as follows: (i) We propose an end-to-end network to learn cross-modal correlation between audio and lyrics, where feature extraction and correlation learning are simultaneously performed and joint representation is learned by considering temporal structures. (ii) And, as for feature extraction, we further represent an audio signal by a short sequence of local summaries (VGG16 features) and apply a recurrent neural network to compute a compact feature that better learns the temporal structures of music audio. Experimental results, using audio to retrieve lyrics or using lyrics to retrieve audio, verify the effectiveness of the proposed deep correlation learning architectures in cross-modal music retrieval. Yi Yu 0001, Suhua Tang, Francisco Raposo 0001, Lei Chen 0002 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2019 | Reducing false wake-up in contention-based wake-up control of wireless LANsabstractThis paper studies the potential problem and performance when tightly integrating a low power wake-up radio (WuR) and a power-hungry wireless LAN (WLAN) module for energy efficient channel access. In this model, a WuR monitors the channel, performs carrier sense, and activates its co-located WLAN module when the channel becomes ready for transmission. Different from previous methods, the node that will be activated is not decided in advance, but decided by distributed contention. Because of the wake-up latency of WLAN modules, multiple nodes may be falsely activated, except the node that will actually transmit. This is called a false wake-up problem and it is solved from three aspects in this work: (i) resetting backoff counter of each node in a way as if it is frozen in a wake-up period, (ii) reducing false wake-up time by immediately putting a WLAN module into sleep once a false wake-up is inferred, and (iii) reducing false wake-up probability by adjusting contention window. Analysis shows that false wake-ups, instead of collisions, become the dominant energy overhead. Extensive simulations confirm that the proposed method (WuR-ESOC) effectively reduces energy overhead, by up to 60% compared with state-of-the-arts, achieving a better tradeoff between throughput and energy consumption. Suhua Tang, Sadao Obana |
Wirel. Networks | 1 |
| 2018 | Efficient Collection of Road and Traffic Information by CCN-based Inter-Vehicle CommunicationsabstractInter-vehicle communication (IVC), which will play an important role in applications such as driving safety support system via periodical exchange of position and speed information among vehicles and autonomous driving which relies on local distribution of road and traffic information, has attracted much attention recently. But IVC does not scale well with the number of requests in the request/response-based communication. In this paper, we propose to apply the technique of Content Centric Network (CCN) to solve this problem, by a cross-layer design, which combines relay vehicle selection, packet forwarding, and cache search, in the same framework. First, we study 1) unique content naming and 2) robust routing (relay vehicle selection, packet forwarding and cache search) to realize the basic function of CCN. Next, we extend the routing method so as to overcome the cache miss problem in the mobile environment. Extensive simulation evaluations confirm the effectiveness of the proposed methods. Compared with the method without caching, the basic method effectively improves system performance. By mitigating the cache miss problem, the extended method (ECCN-IVC) further reduces the average hop count of data packets by up to 46%, and improves the content acquisition success rate by up to 204%. Takanori Nakazawa, Sadao Obana, Suhua Tang |
APCC | 3 |
| 2018 | Deep CNN Denoiser and Multi-layer Neighbor Component Embedding for Face HallucinationabstractMost of the current face hallucination methods, whether they are shallow learning-based or deep learning-based, all try to learn a relationship model between Low-Resolution (LR) and High-Resolution (HR) spaces with the help of a training set. They mainly focus on modeling image prior through either model-based optimization or discriminative inference learning. However, when the input LR face is tiny, the learned prior knowledge is no longer effective and their performance will drop sharply. To solve this problem, in this paper we propose a general face hallucination method that can integrate model-based optimization and discriminative inference. In particular, to exploit the model based prior, the Deep Convolutional Neural Networks (CNN) denoiser prior is plugged into the super-resolution optimization model with the aid of image-adaptive Laplacian regularization. Additionally, we further develop a high-frequency details compensation method by dividing the face image to facial components and performing face hallucination in a multi-layer neighbor embedding manner. Experiments demonstrate that the proposed method can achieve promising super-resolution results for tiny input LR faces. Junjun Jiang, Yi Yu 0001, Suhua Tang, Jiayi Ma 0001 |
IJCAI | 4 |
| 2017 | Compact LBP and WLBP descriptor with magnitude and direction difference for face recognitionabstractIn this paper, we propose a novel descriptor for face recognition on grayscale images, depth images and 2D+depth images. It is a compact and effective descriptor computed from the magnitude and the direction difference. It can be concatenated with conventional descriptors such as well-known Local Binary Pattern (LBP) and Weber Local Binary Pattern (WLBP), to enhance their discrimination capability. To evaluate the performance of our descriptor, we conducted extensive experiments on three types of images using four different databases. The experimental results demonstrate the robustness and superiority of our approach, and the performances of our new descriptor surpass that without magnitude and direction difference. At the end, we further compare our descriptor with Convolution Neural Network (CNN) to show the compactness and effectiveness of the proposed approach. Soo-Chang Pei, Mei-Shuo Chen, Yi Yu 0001, Suhua Tang, Chunlin Zhong |
ICIP | 4 |
| 2017 | Context-patch based face hallucination via thresholding locality-constrained representation and reproducing learningabstractFace hallucination, which refers to predicting a HighResolution (HR) face image from an observed Low-Resolution (LR) one, is a challenging problem. Most state-of-the-arts employ local face structure prior to estimate the optimal representations for each patch by the training patches of the same position, and achieve good reconstruction performance. However, they do not take into account the contextual information of image patch, which is very useful for the expression of human face. Different from position-patch based methods, in this paper we leverage the contextual information and develop a robust and efficient context-patch face hallucination algorithm, called Thresholding Locality-constrained Representation with Reproducing learning (TLcR-RL). In TLcR-RL, we use a thresholding strategy to enhance the stability of patch representation and the reconstruction accuracy. Additionally, we develop a reproducing learning to iteratively enhance the estimated result by adding the estimated HR face to the training set. Experiments demonstrate that the performance of our proposed framework has a substantial increase when compared to state-of-the-arts, including recently proposed deep learning based method. Junjun Jiang, Yi Yu 0001, Suhua Tang, Jiayi Ma 0001, Guo-Jun Qi, Akiko Aizawa |
ICME | 3 |
| 2017 | VenueNet: Fine-Grained Venue Discovery by Deep Correlation LearningabstractVenue photos, as a new type of multimedia contents, are exploding on the Internet because users like to take photos and share with their friends in which venue they spent time and what impressed them there. Discovering a venue by a social photo is very useful for supplementing venue retrieval and recommendation. However, little research focused on fine-grained venue discovery by leveraging multimodal venue dataset. In this paper, we present the first multimodal dataset specially built for venue discovery, which includes venue photos, descriptions, and categories. Using this dataset, we propose a novel framework for fine-grained venue discovery through correlating venue photos and descriptions, aiming to learn a VenueNet representing a knowledge base and association for venues and their properties in different modalities. In the training phase, visual and textual features of the same venues, by two sub-networks, are respectively mapped to a same semantic space, in which canonical correlation analysis (CCA) is applied to these features to train the two sub-networks. In the query phase, given a photo, its correlation with textual features in the dataset is analyzed to find the most similar venue. Experimental results verify the practicability of the Deep CCA model for fine-grained venue discovery from large-scale multimodal dataset. Yi Yu 0001, Suhua Tang, Kiyoharu Aizawa, Akiko Aizawa |
ISM | 2 |
| 2017 | Sink-Based Centralized Transmission Scheduling by Using Asymmetric Communication and Wake-Up RadioabstractIn large-scale wireless sensor networks (WSNs) for wild environments, nodes once deployed run only on battery. Moreover, nodes close to a sink consume energy more quickly than other nodes due to packet forwarding. Mobile sink is a good solution to this issue, although it causes two new problems to nodes: (i) overhead of updating routing information and (ii) increased operating time due to aperiodic query. To solve these problems, this paper proposes an energy efficient data collection protocol for mobile sink, where sink-based centralized transmission scheduling (SC-Sched) is realized by using asymmetric communication and wake-up radio (WuR). Specifically, nodes do not update routing information. Instead, the sink determines the transmission order. In addition, each node is equipped with a WuR. At the time that a packet transmission between two nodes is scheduled, the sink transmits a wake-up message using a large transmission power, directly activating both nodes simultaneously to communicate. Extensive simulation evaluations confirm that the proposed method is more energy efficient than conventional methods, reducing energy consumption to 1/9 in a WSN with 600 nodes. This method is further enhanced to deal with frame loss caused by multipath fading, improving frame delivery rate meanwhile suppressing energy consumption and data collection time. Masanari Iwata, Suhua Tang, Sadao Obana |
WCNC | 2 |
| 2017 | Energy Efficient Downlink Transmission in Wireless LANs by Using Low-Power Wake-Up RadioabstractIn the downlink of a wireless LAN, power-save mode is a typical method to reduce power consumption. However, it usually causes large delay. Recently, remote wake-up control via a low-power wake-up radio (WuR) has been introduced to activate a node to instantly receive packets from an access point (AP). But link quality is not taken into account and protocol overhead of wake-up per node is relatively large. To solve these problems, in this paper, a broadcast-based wake-up control framework is proposed, and a low-power WuR is used to receive traffic indication map from an AP, monitor link quality, and perform carrier sense. Among the nodes which have packets buffered at the AP, only those whose SNR is above a threshold will be activated, contending via a proper contention window to receive packets from the AP. Optimal SNR threshold, deduced by theoretical analysis, helps to reduce transmission collisions and false wake-ups (caused by wake-up latency) and improve transmission rate. Extensive simulations confirm that the proposed method (i) effectively reduces power consumption of nodes compared with other methods, (ii) has less delay than power-save mode in times of light traffic, and (iii) achieves higher throughput than other methods in the saturation state. Suhua Tang, Sadao Obana |
Wirel. Commun. Mob. Comput. | 1 |
| 2016 | Tight Integration of Wake-Up Radio in Wireless LANs and the Impact of Wake-Up LatencyabstractIn wireless LANs (WLANs), a power-hungry transceiver is used for both the control (carrier sense) and data exchanges. As a result, when many nodes contend to access a same channel, the WLAN transceiver performs carrier sense during most time, which wastes much energy. To solve this problem, we study separating the control and data operations in WLANs, by tightly integrating a lower power wake-up radio (WuR) with the WLAN module. Specifically, the WuR is used to monitor the channel and conduct carrier sense. It activates the WLAN module for actual transmissions when the channel gets ready. This is a contention-based self wake-up, which is different from previous schemes where a specific receiver is remotely activated. Due to the hardware constraint, a WLAN module is susceptible to a non-negligible wake-up latency. Our analysis shows that this wake-up latency may break the carrier sense mechanism and lead to false wake-up events. Then, we propose to recover the carrier sense mechanism by resetting the backoff counter of a falsely activated node in a way as if it is frozen in the wake-up period. Simulation evaluations confirm that the proposed scheme effectively mitigates the impact of wake-up latency by reducing the duty time of WLAN modules. Suhua Tang, Sadao Obana |
GLOBECOM | 1 |
| 2016 | PROMPT: Personalized User Tag Recommendation for Social Media Photos Leveraging Personal and Social ContextsabstractSocial media platforms such as Flickr allow users to annotate photos with descriptive keywords, called, tags with the goal of making multimedia content easily understandable, searchable, and discoverable. However, manual annotation is very time-consuming and cumbersome for most users, which makes it difficult to search relevant photos. Moreover, predicted tags for a photo are not necessarily relevant to users' interests. Thus, it necessitates for an automatic tag prediction system that considers users' interests and describes objective aspects of the photo such as visual content and activities. To this end, this paper presents a tag recommendation system, called, PROMPT, that recommends personalized tags for a given photo leveraging personal and social contexts. Specifically, first, we determine a group of users who have similar tagging behavior as the user of the photo, which is very useful in recommending personalized tags. Next, we find candidate tags from visual content, textual metadata, and tags of neighboring photos, and recommends five most suitable tags. We initialize scores of the candidate tags using asymmetric tag co-occurrence probabilities and normalized scores of tags after neighbor voting, and later perform random walk to promote the tags that have many close neighbors and weaken isolated tags. Finally, we recommend top five user tags to the given photo. Experimental results on a Flickr dataset (46,700 photos in the test set and 28 million photos in the train set) with 1,540 unique user tags confirm that the proposed algorithm outperforms state-of-the-arts. Rajiv Ratn Shah, Anupam Samanta, Yi Yu 0001, Suhua Tang, Roger Zimmermann |
ISM | 5 |
| 2016 | Energy and spectrum efficient wireless LAN by tightly integrating low-power wake-up radioabstractCSMA is the de facto standard in wireless LANs (WLANs), where one transceiver is used for both the control and data planes. When many nodes contend to access a same channel, the WLAN transceiver performs carrier sense during most time, which wastes much energy. Dynamic adjustment of parameters (contention window) is suggested in previous works for improving network throughput, but this is also realized at the cost of continuously monitoring the channel with large power consumption. To solve these problems, in this paper, we suggest tightly integrating a lower power wake-up radio (WuR) with the power-hungry WLAN module. Specifically, (i) The WuR is used to monitor the channel and conduct carrier sense. It activates the WLAN module for actual transmissions when the channel gets ready. (ii) The WuR is used to measure the inter-frame space, based on which contention window is adjusted accordingly. Extensive simulation evaluations confirm that the proposed scheme (WuR-CSMA) effectively reduces the duty ratio of WLAN modules and improves system throughput, achieving both energy and spectral efficiency compared with the conventional schemes. Suhua Tang, Chao Zhang 0003, Hiroyuki Yomo, Sadao Obana |
PIMRC | 1 |
| 2016 | Leveraging multimodal information for event summarization and concept-level sentiment analysis
Rajiv Ratn Shah, Yi Yu 0001, Akshay Verma, Suhua Tang, Anwar Dilawar Shaikh, Roger Zimmermann |
Knowl. Based Syst. | 4 |
| 2015 | Exploiting frame length of 802.15.4g signals for wake-up control in sensor networksabstractSensor networks are playing more and more important roles, e.g., accurately monitoring the generation and consumption of electric power for the purpose of smart grid. In these applications, sensor nodes, working with battery, should be put into sleep in the idle state to prolong network lifetime and be activated on demand to ensure real-time response. State-of-the-art schemes merely exploiting duty-cycling cannot simultaneously satisfy these two requirements. In this paper, we exploit frame length of 802.15.4g signals for the wake-up control of sensor nodes. A wake-up ID is modulated onto frame lengths of consecutive signals defined by the mode switch mechanism, transmitted by a sensor node using the standard 802.15.4g protocol, and detected by a non-802.15.4g, low-power wake-up receiver. A prototype wake-up receiver is implemented by using FPGA. It consists of two-stage wake-up control in the sense that duty-cycling is also exploited in the detection of wake-up signals. The proposed scheme achieves a better tradeoff between power-consumption and wake-up latency, compared with the conventional duty-cycling schemes. The experimental evaluation confirms that the wake-up control meets the sensitivity requirement of data communications. Suhua Tang, Hiroyuki Yomo, Shinji Yamaguchi, Akio Hasegawa, Sadao Obana |
WCNC | 1 |
| 2014 | ATLAS: Automatic Temporal Segmentation and Annotation of Lecture Videos Based on Modelling Transition TimeabstractThe number of lecture videos available is increasing rapidly, though there is still insufficient accessibility and traceability of lecture video contents. Specifically, it is very desirable to enable people to navigate and access specific slides or topics within lecture videos. To this end, this paper presents the ATLAS system for the VideoLectures.NET challenge (MediaMixer, transLectures) to automatically perform the temporal segmentation and annotation of lecture videos. ATLAS has two main novelties: (i) a SVMhmm model is proposed to learn temporal transition cues and (ii) a fusion scheme is suggested to combine transition cues extracted from heterogeneous information of lecture videos. According to our initial experiments on videos provided by VideoLectures.NET, the proposed algorithm is able to segment and annotate knowledge structures based on fusing temporal transition cues and the evaluation results are very encouraging, which confirms the effectiveness of our ATLAS system. Rajiv Ratn Shah, Yi Yu 0001, Anwar Dilawar Shaikh, Suhua Tang, Roger Zimmermann |
ACM Multimedia | 4 |
| 2014 | Distributed Multiuser Scheduling for Improving Throughput of Wireless LANabstractIn wireless LANs, the performance of CSMA/CA might be degraded by several problems: (i) severe collisions in the uplink, (ii) head-of-line problem caused by fading in the downlink, and (iii) serious unfairness between uplink and downlink. In this paper, a distributed multiuser scheduling (DMUS) scheme is proposed to simultaneously address these problems. In DMUS, a node (i) computes its normalized SNR (signal to noise ratio) as the ratio of its instantaneous SNR to its average SNR, and (ii) contends via a contention window (CW) for the channel to initiate its uplink or downlink transmission when its normalized SNR is greater than a threshold. The contribution is threefold: (i) All three problems are solved in a unified framework by applying multiuser diversity in both uplink and downlink. Fresh SNR is exploited for distributed scheduling meanwhile airtime fairness is retained. (ii) SNR threshold and CW are jointly optimized to maximize throughput, taking into account time-variant link quality, collision probability and protocol overhead. (iii) Network performance is theoretically analyzed. Extensive simulations confirm that DMUS greatly improves total throughput under almost all scenarios compared with both the contention-based CSMA/CA scheme and the contention-free PCF scheme. Suhua Tang |
IEEE Trans. Wirel. Commun. | 1 |
| 2014 | Iterative receiver for amplify-and-forward relay networks with unknown noise correlationabstractABSTRACT For amplify‐and‐forward relay networks, we propose an iterative scheme to estimate channel and detect information symbols for the multi‐antenna destination in spatially correlated noise. The equivalent channel coefficients and noise covariance are estimated by expectation–maximization algorithm. In addition, we discuss the initialization of iteration and analyze the modified Cramér–Rao bound to show the performance of the proposed iterative estimation. Moreover, on the basis of the structure of the proposed iterative estimator, a joint channel estimation and detection receiver is also provided. Finally, simulation results show that the proposed channel estimator and receiver can achieve the optimal performances in amplify‐and‐forward relay networks with unknown noise correlation. Copyright © 2012 John Wiley & Sons, Ltd. Chao Zhang 0003, Suhua Tang, Pinyi Ren |
Wirel. Commun. Mob. Comput. | 2 |
| 2013 | Edge-based locality sensitive hashing for efficient geo-fencing applicationabstractGeo-fencing is a promising technique for emerging location-based services. Its two basic spatial predicates, INSIDE and WITHIN pairings between points and polygons, can be addressed by state-of-the-art methods such as the crossing number algorithm. In the era of big-data, however, geo-fencing has to process millions of points and hundreds of polygons or even more in real-time. In this paper, we propose an efficient algorithm to improve the scalability of geo-fencing, which consists of two main stages. At the first stage, an R-tree is used to quickly detect whether a point is inside the minimum bounding rectangle of a polygon. In the second stage, instead of an exhaustive search, we design an edge-based locality sensitive hashing scheme adapted to the crossing number algorithm. As for the case of WITHIN detection, a probing scheme is suggested to locate adjacent buckets so as to check all edges near to a target point. By further exploiting batch processing and multi-threading programming, our algorithm can achieve a fast speed while retaining 100% accuracy over all training datasets provided by the GIS Cup 2013 organizers. Yi Yu 0001, Suhua Tang, Roger Zimmermann |
SIGSPATIAL/GIS | 2 |
| 2012 | Receiver design for realizing on-demand WiFi wake-up using WLAN signalsabstractIn this paper, we design a simple, low-cost, and low-power wake-up receiver which can be used for an IEEE 802.11-compliant device to remotely wake up the other devices by utilizing its own wireless LAN (WLAN) signals. A typical usage scenario of such a wake-up receiver is energy management of WiFi device: a device equipped with the wake-up receiver turns WiFi interface off when there is no communication demand, which is powered-on only when the wake-up receiver detects a wake-up signal transmitted by the other WiFi device. The employed wake-up mechanism utilizes the length of 802.11 data frame generated by a WiFi transmitter to differentiate the information conveyed to the wake-up receiver. The wake-up receiver is designed to reliably detect the length of transmitted data frame only with simple envelope detection and limited signal processing. We develop a prototype of the wake-up receiver and investigate the detection performance of the envelope of 802.11 signals. Based on the obtained experimental results, we select appropriate parameters employed by the wake-up receiver to improve the detection performance. Our numerical results show that the proposed wake-up receiver achieves much larger detection range than the off-the-shelf, commercial receiver having the similar functionality. Hiroyuki Yomo, Yoshihisa Kondo, Noboru Miyamoto, Suhua Tang, Masahito Iwai, Tetsuya Ito |
GLOBECOM | 4 |
| 2012 | Exploiting burst transmission and partial correlation for reliable wake-up signaling in Radio-On-Demand WLANsabstractRecent investigations show that (i) access points (APs) of wireless local area networks (WLANs) are idle during much of the time, and, (ii) an AP in its idle state without forwarding any packets still consumes a large percentage of power. Therefore, it is necessary to put idle APs into sleep so as to realize green WLANs. The problem, however, is how to quickly and reliably activate APs from sleep when nodes initiate new data flows. Aiming at realizing Radio-On-Demand WLANs, in this paper, we suggest exploiting burst transmission of WLAN frames to convey wake-up IDs from nodes to APs. Our contribution is two-fold: (i) The burst transmission prevents interfering WLAN signals from breaking in, and, (ii) We re-interpret the sequence of WLAN frames for wake-up signaling as an equivalent ID. Based on the analysis of hamming distance among equivalent IDs, we further suggest using partial correlation to reduce the error rate of wake-up signals. The effectiveness of the proposed scheme is confirmed by both theoretical analysis and simulation evaluations. Suhua Tang, Hiroyuki Yomo, Yoshihisa Kondo, Sadao Obana |
ICC | 1 |
| 2012 | Wake-up ID and protocol design for radio-on-demand wireless LANabstractThis paper investigates wake-up ID and protocol design for radio-on-demand (ROD) wireless LAN (WLAN) where wake-up radio is applied to access point (AP) in order to save energy consumed by WLAN. Each AP in ROD WLAN is transited to a sleep state when there is no associated user. A wake-up receiver installed into each AP is used to detect a wake-up signal transmitted by a station (STA) upon communications demands. Each STA specifies an AP to wake up by embedding ID of the target AP into the wake-up signal. In this paper, we propose a wake-up ID assignment and ID matching which can reduce the probability of false wake-up caused by bit errors over wireless channels. The proposed scheme generates wake-up ID of each AP based on ESSID in such a way that certain number of hamming distances is maintained among different wake-up IDs. The AP wakes up when the hamming distance between the received ID, possibly containing bit errors, and its assigned ID is less than a predetermined value. The numerical results obtained by theoretical analysis and computer simulation show that the proposed scheme can effectively reduce the false wake-up probabilities with short ID length and very simple operations at wake-up receivers. We also propose a wake-up protocol for ROD WLAN to reduce energy wastefully consumed by APs which are redundantly woken up with the proposed ID assignment/matching. Our simulation results show that ROD WLAN with the proposed protocol achieves much better energy-efficiency than ROD WLAN without the proposed protocol and WLAN without applying ROD technologies. Hiroyuki Yomo, Yoshihisa Kondo, Kosuke Namba, Suhua Tang, Takatoshi Kimura, Tetsuya Ito |
PIMRC | 4 |
| 2012 | EM Algorithm Based Channel Estimation for Amplify-and-Forward Relay Networks with Unknown Noise CorrelationabstractDue to common interference or noise propagation, noise correlation between relays could occur in Amplify-and-Forward relay networks. To estimate channel coefficient, the noise covariance is required in traditional channel estimations. On the other hand, we also need channel coefficient to estimate the noise covariance. Therefore,traditional channel estimators can not be utilized when noise correlation is unknown. In this paper, we propose an Expectation-Maximization algorithm based iterative channel estimator to solve this problem. Moreover, we analyze the modified Cram'er-Rao bound to show the performance of the proposed channel estimation. Finally, simulation results show that the proposed channel estimator can work well in Amplify-and- Forward relay networks with unknown noise correlation. Chao Zhang 0003, Suhua Tang, Pinyi Ren |
VTC Fall | 2 |
| 2012 | Energy-efficient WLAN with on-demand AP wake-up using IEEE 802.11 frame length modulation
Yoshihisa Kondo, Hiroyuki Yomo, Suhua Tang, Masahito Iwai, Toshiyasu Tanaka, Hideo Tsutsui, Sadao Obana |
Comput. Commun. | 3 |
| 2011 | Wakeup Receiver for Radio-On-Demand Wireless LANsabstractAccess points (APs) of wireless LANs (WLANs), always powered on and ready to serve mobile nodes, consume a large amount of power in total, but are idle during much of the time. Previous protocol-based sleep-wakeup scheduling schemes partially solve this problem, but the large wakeup delay remains a problem. Aiming at realizing Radio-On-Demand WLANs, in this paper, we suggest using an additional wakeup transceiver to convey wakeup signals from nodes to APs. APs in the sleep mode are activated by nodes when new data flows are initiated, where the wakeup delay is very low. The proposed wakeup transceiver works on the 2.4GHz ISM band and shares antenna with a co-located WLAN module to reduce hardware cost. The wakeup receiver is designed to be simple but very reliable, and consume a very low power. The contribution of this paper is two-fold: (i) Wakeup signals are designed to co-exist with WLAN signals by exploiting the carrier sense mechanism of WLAN devices, and, (ii), explicit signal recognition is used to achieve an extremely low false wakeup probability in an environment where the number of WLAN signals is overwhelming. Testbed experiments confirm that the proposed scheme, besides its simplicity, has good performance in both frame error rate and false wakeup probability. Suhua Tang, Hiroyuki Yomo, Yoshihisa Kondo, Sadao Obana |
GLOBECOM | 1 |
| 2011 | Wake-up radio using IEEE 802.11 frame length modulation for Radio-On-Demand wireless LANabstractIn this paper, we introduce Radio-On-Demand (ROD) wireless LAN (WLAN) in which access points (APs) are put into a sleep mode during idle periods and woken up by stations (STAs) upon communications demands. The on-demand wake-up is realized by a wake-up receiver which is equipped with each AP and is used to detect a wake-up signal transmitted by STA. In this paper, in order to reduce the hardware installation cost at STA, we advocate to utilize wireless LAN frames transmitted by each STA as a wake-up signal to awake the target AP. The STA generates a wake-up signal by devising WLAN signal: each STA creates a series of WLAN frames with different length to which the information on wake-up ID is embedded. The wake-up receiver extracts the wake-up ID from the received frames with a simple detector which ensures its low-power operation. We evaluate false negative (STA fails to wake up the target AP) and false positive (AP falsely wakes up without an intended wake-up signal) probabilities of our proposed on-demand wake-up scheme with computer simulations. The numerical results show that the proposed scheme achieves the false negative probability of about 10-2when the detection error ratio of `1' is less than 10-3. We also show that the false positive probability can be largely reduced by employing long WLAN frames to generate each wake-up signal. These results confirm that the proposed wake-up scheme is a promising approach to reducing wasteful energy consumed by idle APs in WLAN. Yoshihisa Kondo, Hiroyuki Yomo, Suhua Tang, Masahito Iwai, Toshiyasu Tanaka, Hideo Tsutsui, Sadao Obana |
PIMRC | 3 |
| 2011 | Improving throughput of wireless LANs with transmit power control and slotted channel accessabstractThroughput of wireless LANs is interference-limited. Newly deployed access points (APs) can hardly bring throughput gain without careful design. Although transmit power control (TPC) can improve spatial reuse of the channels, the achievable throughput also heavily depends on other factors, such as rate adaptation, which are seldom touched in previous works. In addition, TPC protocols, designed with the assumption that all nodes have the same capability, do not work well when legacy nodes not supporting TPC coexist. In this paper, (i) we exploit joint rate adaptation and TPC to improve system performance and suggest choosing power to maximize throughput after taking tradeoff between transmit rate and spatial reuse of channels, and, (ii) we suggest using the NAV (network allocation vector) mechanism to create virtual slots in the scenario where legacy nodes coexist with new nodes running TPC. Simulation results show that (i) the proposed TPC scheme can effectively improve both total throughput and fairness, and, (ii) the slotted channel access scheme ensures that throughput gain can be achieved by nodes running TPC while the performance of legacy nodes not supporting TPC is not degraded. Suhua Tang, Akio Hasegawa, Riichiro Nagareda, Akito Kitaura, Tatsuo Shibata, Sadao Obana |
PIMRC | 1 |
| 2010 | Achieving Full Rate Network Coding with Constellation Compatible Modulation and CodingabstractNetwork coding is an effective method to improving relay efficiency by reducing the number of transmissions. However, its performance is limited by several factors such as packet length mismatch and rate mismatch. Although the former may be solved by re-framing, the latter remains a challenge and is likely to greatly degrade the efficiency of network coding. In this paper, we re-interpret network coding as a mapping of modulation constellation. On this basis, we extend such mapping to enable simultaneous use of different modulations by nesting the low-level constellation as a subset of the high level constellation. When relay links have different qualities, the messages of different flows are combined together in such a way that for each relay link its desired message is transmitted at its own highest rate. Compared with previous solutions to rate mismatch, the proposed scheme achieves the full rate of all relay links on the broadcast channel. Suhua Tang, Hiroyuki Yomo, Tetsuro Ueda, Ryu Miura, Sadao Obana |
GLOBECOM | 1 |
| 2009 | Distributed Multi-User Scheduling for Improving Throughput of Wireless LANabstractCarrier sense multi-access (CSMA) is a typical method to share the common channel in a wireless LAN (WLAN). It works fairly well in times of light traffic. However as the number of nodes in a WLAN increases quickly, severe collision greatly degrades network performance. In this paper we propose a distributed multi-user scheduling (DMUS) scheme to solve this problem, taking time-variant link quality and rate adaptation into account. Instead of all nodes, only nodes with high instantaneous link quality are allowed to contend for the channel. By setting a suitable SNR threshold, at any instance only a small percentage of nodes join the contention. As a result, collision is mitigated, fairness is retained by independent fading, and the total throughput is increased since transmissions are finished at higher rates. Simulation results show when there are 40 nodes in a WLAN, the DMUS scheme improves total throughput by up to 49.6% compared with the contention-free scheme, and by up to 194.6% compared with the CSMA/CA scheme. Suhua Tang, Ryu Miura, Sadao Obana |
ICC | 1 |
| 2009 | Exploiting network coding for pseudo bidirectional relay in wireless LANabstractNetwork coding is an effective method to improve forwarding efficiency in multi-hop wireless networks. Previous network coding schemes proposed for unicast transmissions focus on scheduling where paths are already established, and assume a fixed rate for all links, neglecting link rate heterogeneity. In this paper we study a network coding based pseudo bidirectional relay scheme for wireless LAN, taking link quality into account. We first show when network coding based relay improves channel efficiency and how to select such a relay in a distributed way. Then we enhance network coding based transmissions from three aspects: (i) Coordinated rate adaptation is adopted to reduce the probability with which a priori packet is missing in times of network decoding. (ii) The retransmission scheduling completely solves the lack-of-a-priori problem. (iii) Piggybacking receiving status in ACK/NAK further enables opportunistic reception and avoids unnecessary transmissions. Simulation results confirm that compared with direct transmissions, the proposed relay scheme can reduce per-packet transmission time by up to 39% in Rayleigh fading environment. Suhua Tang, Hiroyuki Yomo, Mehdad N. Shirazi, Tetsuro Ueda, Ryu Miura, Sadao Obana |
PIMRC | 1 |
| 2009 | Message Dissemination in Inter-Vehicle CDMA Networks for Safety Driving SupportabstractAlthough the near-far effect has been considered to be the major issue preventing CDMA from being used in ad-hoc networks, in this paper, we show that the near-far effect is not a severe issue in inter-vehicle networks for safety driving support, where packet transmissions are performed in the broadcast manner. Indeed, the near-far effect provides extremely reliable transmissions between near nodes, regardless of node density, which can not be achieved by CSMA/CA. However, CDMA can not be directly applied in realistic traffic accident scenarios, where highly reliable transmissions are required between far nodes as well. This paper proposes to apply packet forwarding and transmission scheduling methods that try to expand the area, where reliable transmissions are achievable. Simulation results show that the proposed scheme achieves approximately 100% of delivery ratio and 4 milliseconds of delay in a realistic traffic accident scenario, where CSMA/CA achieves approximately 60% of delivery ratio and 80 milliseconds of delay. Oyunchimeg Shagdar, Takashi Ohyama, Mehdad N. Shirazi, Suhua Tang, Ryutaro Suzuki, Ryu Miura, Sadao Obana |
VTC Spring | 4 |
| 2008 | Opportunistic cooperation and selective forwarding, a virtual MIMO scheme for wireless networksabstractMultipath fading greatly degrades system performance of wireless networks. Conventionally diversity is introduced to mitigate fading. In this paper we study space diversity in multi-hop wireless networks under power constraints. Point-to-point links are extended to group-to-group virtual MIMO links. Then we propose a cross-layer design of MAC and routing that exploits opportunistic cooperation and selective forwarding (OCSF). Packets are forwarded around the pre-calculated anchor route and the proposed OCSF scheme enables local post selection of both transmitter and forwarder according to channel state information. Simulation results confirm that OCSF outperforms existing schemes that utilize distributed space-time coding or selection diversity. Suhua Tang, Oyunchimeg Shagdar, Mehdad N. Shirazi, Ryutaro Suzuki, Sadao Obana |
PIMRC | 1 |
| 2007 | An Opportunistic Progressive Routing (OPR) Protocol Maximizing Channel EfficiencyabstractIn this paper we study the channel efficiency of mobile ad hoc networks (MANET) and try to trade off between the progress and the instantaneous rate when forwarding packets in the fading environment. We define a new metric-bit transfer speed (BTS)-as the ratio of the progress made towards the destination to the equivalent time taken to transfer a payload bit. This metric takes the overhead, rate and progress into account. Then we propose an opportunistic progressive routing (OPR) protocol, jointly optimizing routing and MAC by the cross-layer design. In OPR, a node selects the forwarder with the biggest BTS for a packet and the forwarder changes as the packet size or channel status varies. The extensive simulation shows that OPR greatly reduces channel occupation time and packet loss compared with the normalized advance (NADV) scheme and contention-based forwarding (CBF) scheme. Suhua Tang, Ryutaro Suzuki, Sadao Obana |
GLOBECOM | 1 |
| 2007 | Improving Routing Performance Under the Fading Environment by Utilizing Position InformationabstractPerformance of mobile ad hoc networks (MANET) routing protocols may be greatly degraded due to involvement of weak links in the routes. One solution is to use directional antenna to improve link quality. The other is to less prefer weak links. In the experiments we found that even with the two methods the system performance is still limited in the presence of mobility and multipath fading due to the following facts: (1) beam scanning overhead and frequent beam variations, (2) route instability due to metric variation and false link breaks. Then we improve it by three schemes: (i) Calculate the antenna beam by position information, (ii) Calculate a stable link metric from the expectation value of received signal strength indication (RSSI). (iii) Avoid false link breaks. These schemes are suitable for the scenarios where both line of sight path and multipath fading exist. The experiment and simulation results indicate that the schemes can effectively reduces PER and improve throughput of AODV in the Rician fading situations. Suhua Tang, Masahiro Watanabe, Naoto Kadowaki, Sadao Obana |
WCNC | 1 |
| 2006 | Improving Performance of On-demand Routing under Multipath FadingabstractDespite many routing protocols proposed for mobile ad hoc networks few have considered link quality and route stability. Weak links contained in the routes often lead to route breaks. Multipath fading and mobility make link quality timevariant. When the link metric is associated with link quality the routes may change frequently. In times of transmission failure due to collision or fading, the routing packets are susceptible to loss owing to the lack of retransmission mechanism for the broadcast packets. For these reasons, the routes become unstable. In this paper, we investigate the long-term link quality and route stability in the on-demand routing protocols. We compare the link metric policies, optimize the route discovery and use the soft detection scheme to avoid false link breaks. The simulation results show that both Packet Error Ratio (PER) and throughput in the enhanced on-demand routing protocol can be greatly improved under the Rayleigh fading environment. Suhua Tang, Masahiro Watanabe, Naoto Kadowaki, Sadao Obana |
VTC Spring | 1 |
| 2005 | Trigger Update based Local Optimization for On-demand Routing ProtocolsabstractIn mobile ad hoc networks on-demand routing protocols are very attractive due to low overhead. In these protocols, however, (1) the routes tend to contain weak links that have a low rate and high packet error ratio; (2) the routes are susceptible to breaks in the presence of mobility. Our major concern in this paper is to locally optimize the routes in the on-demand routing protocols in terms of link quality and mobility. The local optimization is realized by two main techniques: relating link metric with received signal strength and preferring the route with a shorter metric through trigger update (TU). In this manner the initial routes converge to the local optimum in the static case, and adapt to topology variations and link quality changes caused by mobility. Also, the routing overhead is controlled. Simulation results show that AODV with the application of local optimization (AODV-TU) achieves a much higher performance than the original AODV Suhua Tang, Masahiro Watanabe, Naoto Kadowaki, Sadao Obana |
LCN | 1 |