Search results for "Speech recognition"
showing 10 items of 357 documents
Cumulative-Sum-Based Localization of Sound Events in Low-Cost Wireless Acoustic Sensor Networks
2014
Wireless acoustic sensor networks (WASNs) are known for their potential applications in multiple areas, such as audio-based surveillance, binaural hearing aids or advanced acoustic monitoring. The knowledge of the spatial position of a source of interest is usually a requirement for many of these applications. Therefore, source localization is an important problem to be addressed in WASNs. Unfortunately, most localization algorithms need costly signal processing stages that prevent them from being implemented in low-cost sensor networks, requiring additional modules for signal acquisition and processing. This paper presents a low-complexity method for acoustic event detection and localizati…
An integrated dialect analysis tool using phonetics and acoustics
2019
This study aimed to verify a computational phonetic and acoustic analysis tool created in the MATLAB environment. A dataset was obtained containing 3 broad American dialects (Northern, Western and New England) from the TIMIT database using words that also appeared in the Swadesh list. Each dialect consisted of 20 speakers uttering 10 sentences. Verification using phonetic comparisons between dialects was made by calculating the Levenshtein distance in Gabmap and the proposed software tool. Agreement between the linguistic distances using each analysis method was found. Each tool showed increasing linguistic distance as a function of increasing geographic distance, in a similar shape to Segu…
Usage of HMM-Based Speech Recognition Methods for Automated Determination of a Similarity Level Between Languages
2019
The problem of automated determination of language similarity (or even defining of a distance on the space of languages) could be solved in different ways – working with phonetic transcriptions, with speech recordings or both of them. For the recordings, we propose and test a HMM-based one: in the first part of our article we successfully try language detection, afterwards we are trying to calculate distances between HMM-based models, using different metrics and divergences. The Kullback-Leibler divergence is the only one we got good results with – it means that the calculated distances between languages correspond to analytical understanding of similarity between them. Even if it does not …
Using privacy-transformed speech in the automatic speech recognition acoustic model training
2020
Automatic Speech Recognition (ASR) requires huge amounts of real user speech data to reach state-of-the-art performance. However, speech data conveys sensitive speaker attributes like identity that can be inferred and exploited for malicious purposes. Therefore, there is an interest in the collection of anonymized speech data that is processed by some voice conversion method. In this paper, we evaluate one of the voice conversion methods on Latvian speech data and also investigate if privacy-transformed data can be used to improve ASR acoustic models. Results show the effectiveness of voice conversion against state-of-the-art speaker verification models on Latvian speech and the effectivene…
Speech Activity Detection under Adverse Noisy Conditions at Low SNRs
2021
Speech originating from the noisy environments degrades the speech quality and intelligibility, thus reducing the human perceived Quality of Experience (QoE). For example, surveillance using drone during natural catastrophe needs an efficient speech recognition device to recognise the speech of the frozen human in presence of drone noise to save their life. Therefore, it often requires to pre-process the noisy speech in order to reduce the noise artifacts and enhance the speech. This paper detects the speech activity using Voice Activity Detection (VAD). The VAD distinguishes speech activity (speech presence) and speech inactivity (silence/noise) by extracting the speech features and compar…
Lexical and sublexical units in speech perception.
2009
Saffran, Newport, and Aslin (1996a) found that human infants are sensitive to statistical regularities corresponding to lexical units when hearing an artificial spoken language. Two sorts of segmentation strategies have been proposed to account for this early word-segmentation ability: bracketing strategies, in which infants are assumed to insert boundaries into continuous speech, and clustering strategies, in which infants are assumed to group certain speech sequences together into units (Swingley, 2005). In the present study, we test the predictions of two computational models instantiating each of these strategies i.e., Serial Recurrent Networks: Elman, 1990; and Parser: Perruchet & Vint…
Interaction in Spoken Word Recognition Models: Feedback Helps
2018
Human perception, cognition, and action requires fast integration of bottom-up signals with top-down knowledge and context. A key theoretical perspective in cognitive science is the interactive activation hypothesis: forward and backward flow in bidirectionally connected neural networks allows humans and other biological systems to approximate optimal integration of bottom-up and top-down information under real-world constraints. An alternative view is that online feedback is neither necessary nor helpful; purely feed forward alternatives can be constructed for any feedback system, and online feedback could not improve processing and would preclude veridical perception. In the domain of spo…
Reading traffic signs while driving: Are linguistic word properties relevant in a complex, dynamic environment?
2019
When driving a vehicle, do we read the words displayed on traffic signs just as we do in more standard conditions? In the driving context, stimulus quality is generally worse, and reading has to be performed at the same time as we are doing other tasks. In the present work, we examined the effects of word frequency and word length on reading in such circumstances. A stimulus presentation mimicking the approach to the traffic sign increased the effect of word frequency, but not the effect of word length, on reading latency. In addition, performing the reading task while driving along a simulated route produced similar results. Therefore, in the context of the driving activity, the advantage …
Author response: Individual differences in selective attention predict speech identification at a cocktail party
2016
Spatiotemporal Dynamics of the Processing of Spoken Inflected and Derived Words:A Combined EEG and MEG Study
2011
The spatiotemporal dynamics of the neural processing of spoken morphologically complex words are still an open issue. In the current study, we investigated the time course and neural sources of spoken inflected and derived words using simultaneously recorded electroencephalography (EEG) and magnetoencephalography (MEG) responses. Ten participants (native speakers) listened to inflected, derived, and monomorphemic Finnish words and judged their acceptability. EEG and MEG responses were time-locked to both the stimulus onset and the critical point (suffix onset for complex words, uniqueness point for monomorphemic words). The ERP results showed that inflected words elicited a larger left-late…