Acoustical Science and Technology
Online ISSN : 1347-5177
Print ISSN : 1346-3969
ISSN-L : 0369-4232
Current issue
Displaying 1-11 of 11 articles from this issue
PAPERS
  • Chiho Haruta, Nobutaka Ono
    2026Volume 47Issue 4 Pages 303-313
    Published: July 01, 2026
    Released on J-STAGE: July 01, 2026
    Advance online publication: February 14, 2026
    JOURNAL OPEN ACCESS

    In this paper, we propose an element selection approach for speech enhancement using deep neural networks (DNNs) targeting small devices with complexity constraints, such as hearing aids. Element selection reduces the input dimensionality by selecting specific elements from the input vector. Unlike other dimensionality reduction algorithms such as principal component analysis, element selection does not require multiplications, making it suitable for low-complexity environments. To optimize which elements are selected, we propose two methods: 1) a linear-regression-based method minimizing the regression error in estimating a target vector from the dimensionality-reduced vector and 2) a pruning-based method that selects elements corresponding to the remaining weight coefficients in the first layer after applying structured pruning to a DNN. We evaluate their performance in a speech enhancement task under complexity constraints, assuming a simple fully-connected network, with no more than 4×105 multiplications per inference and an algorithmic delay below 8 ms. Experiments show that the proposed approach under the complexity constraints achieves a scale-invariant source-to-distortion ratio (SI-SDR) improvement of 5.6 dB on average compared to non-processed noisy speech at signal-to-noise ratios −5, 0, and 5 dB, and 2.54 dB SI-SDR improvement compared to simply using only the latest frames.

    Download PDF (848K)
  • Yuto Otani, Shun Sawada, Hidefumi Ohmura, Kouichi Katsurada
    2026Volume 47Issue 4 Pages 314-323
    Published: July 01, 2026
    Released on J-STAGE: July 01, 2026
    Advance online publication: April 07, 2026
    JOURNAL OPEN ACCESS

    This paper presents a novel approach for speech synthesis using articulatory movements captured by real-time magnetic resonance imaging (rtMRI), focusing on fundamental frequency (F0) estimation mechanisms. Although recent rtMRI-based methods have achieved promising results, it remains unclear how F0 information is reproduced, given rtMRI's limited ability to capture vocal fold vibrations. To address this gap, we propose a speech synthesis method that processes only four consecutive rtMRI frames (~150 ms)—preventing reliance on extended linguistic context to infer F0. Our method employs an EfficientNetV2-BiLSTM network that enables sophisticated F0-related feature extraction for mel-spectrogram estimation, followed by a HiFi-GAN vocoder for high-fidelity waveform generation. Evaluations on the ATR 503 sentences rtMRI database demonstrate intelligible speech synthesis with accurate F0 reproduction. Building on these results, we further estimate F0 from single MRI frames, confirming that F0 can be derived without temporal context. To explore the underlying basis, we apply optical flow analysis to visualize subtle articulatory differences associated with F0 control, primarily revealing upward/forward larynx and tongue shifts with increasing F0. Additionally, distinct patterns were observed in male speakers at low F0 ranges. These findings empirically validate the relationship between articulatory configurations and F0 control, demonstrating feasibility in rtMRI-based speech synthesis.

    Download PDF (1119K)
  • Kazunori Harada, Yasuhiro Hiraguri, Takuya Oshima, Yoshinori Saito, Sa ...
    2026Volume 47Issue 4 Pages 324-333
    Published: July 01, 2026
    Released on J-STAGE: July 01, 2026
    Advance online publication: March 20, 2026
    JOURNAL OPEN ACCESS

    This study assesses the health impacts of road traffic noise in Osaka City (Tennoji Ward) and Higashiosaka City, Osaka Prefecture, using strategic noise maps. The ASJ RTN-Model 2018 was employed to estimate noise levels, and three noise reduction scenarios were analyzed: reducing light vehicle power levels, heavy vehicle level and installing porous pavement. The study also combined estimations of building occupancy from open data and noise exposure level, enabling assessments of noise health impact. Results of three scenarios and reference showed that approximately 10% and 5% of the population in Tennoji Ward and Higashiosaka City, respectively, are estimated to be suffering from high annoyance due to traffic noise. While noise reduction measures effectively decreased exposure levels, their impact on ischaemic heart disease incidence was limited, suggesting the need for more comprehensive mitigation strategies. The effectiveness of different scenarios varied between cities, highlighting the importance of considering local urban characteristics in noise management planning. This study demonstrates the practical application of strategic noise mapping for health impact assessment in Japanese cities and provides valuable insights for evidence-based urban noise policy development.

    Download PDF (1508K)
  • Mariko Tsuruta-Hamamura, Hiroshi Hasegawa, Shin-ichiro Iwamiya
    2026Volume 47Issue 4 Pages 334-343
    Published: July 01, 2026
    Released on J-STAGE: July 01, 2026
    Advance online publication: April 08, 2026
    JOURNAL OPEN ACCESS

    Previous studies have reported gender differences in perceived loudness. For instance, women tend to assign higher loudness scores to sounds with the same sound pressure level than men when verbal expressions such as "soft" and "loud" are used to evaluate perceived loudness. However, when a ratio scale was used, gender differences were observed under limited conditions in Chinese participants but not Japanese participants. In this study, to clarify the factors affecting gender differences in loudness perception, we conducted four experiments involving magnitude estimation and magnitude production methods in Japanese participants. We examined gender difference in perceived loudness with respect to changes in sound pressure level. The power exponent α in Stevens' power law, estimated from the experimental results of the four experiments, did not show a clear gender-based difference. According to our results, gender differences in judgment criteria using verbal expression such as "soft" and "loud" might be a principal factor causing gender differences in the evaluation of perceived loudness, at least among Japanese participants.

    Download PDF (661K)
  • Mizuki Iwagami, Yuting Geng, Masato Nakayama, Takanobu Nishiura
    2026Volume 47Issue 4 Pages 344-352
    Published: July 01, 2026
    Released on J-STAGE: July 01, 2026
    Advance online publication: March 06, 2026
    JOURNAL OPEN ACCESS

    Parametric array loudspeakers achieve sharp directivity in audible sound by utilizing nonlinear interactions among ultrasounds in air. Conventionally, pin-spot audio has been realized by emitting ultrasounds separately. However, nonlinear interactions among sideband components lead to speech leakage outside the audio spot. A previous study applied subband decomposition to the sideband of an amplitude-modulated signal, which altered the spectrum of the leaked sound but provided limited controllability because only one sideband was processed. This study proposes a pin-spot audio design that combines double sideband modulation with suppressed carrier and subband decomposition applied across both upper and lower sidebands. By designing the structures of the two sideband spectra, the proposed method controls nonlinear interactions across them, producing more complex patterns in the resulting spectrum in air and reducing speech leakage. Moreover, a logarithmic subband decomposition approximately consistent with perceptual frequency spacing and an asymmetric sideband assignment between the two sidebands are introduced. As a result, speech leakage is reduced not merely by lowering the sound pressure level, but by altering the structure of the demodulated sound spectrum.

    Download PDF (1103K)
TECHNICAL REPORT
  • Shigeaki Amano, Kimiko Yamakawa, Arkadiusz Rojczyk
    2026Volume 47Issue 4 Pages 353-357
    Published: July 01, 2026
    Released on J-STAGE: July 01, 2026
    Advance online publication: April 02, 2026
    JOURNAL OPEN ACCESS

    Previous studies on geminate and singleton consonants have employed the ratio of geminate duration to singleton duration (the GS ratio) as an invariant parameter to nullify the effects of speaking rate variation. However, the validity of the implicit assumption that the GS ratio effectively compensates for duration variations induced by speaking rate has not yet been empirically tested. This study formalized this GS ratio assumption mathematically in two scenarios: linear and logarithmic scales of duration. Furthermore, it examined the validity of this assumption using parameters derived from previous research. Our analysis identified the specific mathematical conditions that must be satisfied for the implicit assumption to hold. The empirical test revealed that these conditions were not met, indicating that the GS ratio varies across different speaking rates. These results suggest that the implicit assumption of the GS ratio is not supported in either linear or logarithmic scales. This contradicts the assumptions in previous studies and indicates that careful verification is necessary when employing the GS ratio.

    Download PDF (384K)
ACOUSTICAL LETTERS
  • Yutao Zhang, Shiori Totsuka, Yuting Geng, Masato Nakayama, Ryo Akama, ...
    2026Volume 47Issue 4 Pages 358-362
    Published: July 01, 2026
    Released on J-STAGE: July 01, 2026
    Advance online publication: April 07, 2026
    JOURNAL OPEN ACCESS

    This letter proposes a new Kuzushiji transcription framework that integrates optical character recognition (OCR) with read-speech automatic speech recognition (ASR) via hiragana-level fusion, without requiring additional model training. The framework uses the transcriber's read-speech as an additional modality to guide beam-search OCR hypothesis selection for Kuzushiji transcription. Each OCR candidate is scored based on its phonetic similarity to the ASR output of the corresponding Kuzushiji read-speech at the hiragana-sequence level. Evaluation results show the effectiveness of the proposed framework in reducing the character error rate in contrast to conventional OCR-only Kuzushiji transcription.

    Download PDF (338K)
  • Kento Hara, Tsuguto Hoshino, Motoki Yairi, Takashi Takeuchi, Philip A. ...
    2026Volume 47Issue 4 Pages 363-367
    Published: July 01, 2026
    Released on J-STAGE: July 01, 2026
    Advance online publication: March 31, 2026
    JOURNAL OPEN ACCESS

    The theory of the Optimal Source Distribution (OSD) has been proposed as the basis for binaural synthesis over loudspeakers. In applying this theory to practical systems, discrete linear loudspeaker arrangements and frequency-band division filtering cause an increase in the condition number of the transfer function matrix. This paper proposes a transfer function matrix reconstruction method by applying gain and delay parameters that are optimized using numerical optimization to reduce the condition number. The effectiveness of the proposed method is experimentally validated using a multiway loudspeaker system based on the OSD principle.

    Download PDF (1255K)
  • Yuta Goshima, Yoichi Haneda
    2026Volume 47Issue 4 Pages 368-371
    Published: July 01, 2026
    Released on J-STAGE: July 01, 2026
    Advance online publication: March 28, 2026
    JOURNAL OPEN ACCESS

    Synthesizing virtual sound sources that traverse a linear loudspeaker array poses a challenge for wave field synthesis (WFS) due to singularities. To address this issue, we propose a time-domain representation of WFS based on the spatial-shifting filter, derived via analytical inverse temporal and spatial Fourier transforms. By applying the stationary phase approximation, the proposed method derives the driving function suitable for sample-by-sample processing, thereby eliminating frame-length latency. The validity of the proposed method is demonstrated through numerical simulations.

    Download PDF (648K)
  • Masahiro Toyoda
    2026Volume 47Issue 4 Pages 372-375
    Published: July 01, 2026
    Released on J-STAGE: July 01, 2026
    Advance online publication: April 01, 2026
    JOURNAL OPEN ACCESS

    The receiving points used for measuring floor impact sound levels are specified as follows: "Within the receiving room, distribute four or more measurement points evenly, each separated by at least 70 cm, with spaces at least 50 cm away from the ceiling, surrounding walls, and floor surface." The energy-averaged sound level over these points is used for evaluation. The present letter verifies whether this averaged value sufficiently represents the floor impact sound levels throughout the receiving room. Using the finite-difference time-domain method, a two-story concrete building was analyzed. The floor impact sound levels were compared between a case where multiple receiving points were installed and a case where only the receiving points specified by the standard were installed. The comparison confirmed that while differences exceeding 5 dB were observed in several frequency bands, the specified points sufficiently represent the floor impact sound level with adequate accuracy across a broad frequency range, including the decisive frequency.

    Download PDF (889K)
feedback
Top