Search results for "Speech recognition"

showing 10 items of 357 documents

Cumulative-Sum-Based Localization of Sound Events in Low-Cost Wireless Acoustic Sensor Networks

2014

Wireless acoustic sensor networks (WASNs) are known for their potential applications in multiple areas, such as audio-based surveillance, binaural hearing aids or advanced acoustic monitoring. The knowledge of the spatial position of a source of interest is usually a requirement for many of these applications. Therefore, source localization is an important problem to be addressed in WASNs. Unfortunately, most localization algorithms need costly signal processing stages that prevent them from being implemented in low-cost sensor networks, requiring additional modules for signal acquisition and processing. This paper presents a low-complexity method for acoustic event detection and localizati…

Sound localizationSignal processingAcoustics and UltrasonicsComputer sciencebusiness.industrySpeech recognitionNode (networking)Real-time computingCUSUMComputational MathematicsSoftware deploymentComputer Science (miscellaneous)WirelessElectrical and Electronic EngineeringbusinessWireless sensor networkChange detectionIEEE/ACM Transactions on Audio, Speech, and Language Processing
researchProduct

An integrated dialect analysis tool using phonetics and acoustics

2019

This study aimed to verify a computational phonetic and acoustic analysis tool created in the MATLAB environment. A dataset was obtained containing 3 broad American dialects (Northern, Western and New England) from the TIMIT database using words that also appeared in the Swadesh list. Each dialect consisted of 20 speakers uttering 10 sentences. Verification using phonetic comparisons between dialects was made by calculating the Levenshtein distance in Gabmap and the proposed software tool. Agreement between the linguistic distances using each analysis method was found. Each tool showed increasing linguistic distance as a function of increasing geographic distance, in a similar shape to Segu…

Space (punctuation)Dialectometry050101 languages & linguisticsLinguistics and LanguageSpeech recognition05 social sciencesPhoneticsLinguistic distanceLevenshtein distance050105 experimental psychologyLanguage and LinguisticsVariation (linguistics)Swadesh listVowel0501 psychology and cognitive sciencesMathematicsLingua
researchProduct

Usage of HMM-Based Speech Recognition Methods for Automated Determination of a Similarity Level Between Languages

2019

The problem of automated determination of language similarity (or even defining of a distance on the space of languages) could be solved in different ways – working with phonetic transcriptions, with speech recordings or both of them. For the recordings, we propose and test a HMM-based one: in the first part of our article we successfully try language detection, afterwards we are trying to calculate distances between HMM-based models, using different metrics and divergences. The Kullback-Leibler divergence is the only one we got good results with – it means that the calculated distances between languages correspond to analytical understanding of similarity between them. Even if it does not …

Space (punctuation)Kullback–Leibler divergenceLanguage identificationSimilarity (network science)Computer scienceSpeech recognitionComputer Science::Computation and Language (Computational Linguistics and Natural Language and Speech Processing)Hidden Markov modelUSableDivergence (statistics)
researchProduct

Using privacy-transformed speech in the automatic speech recognition acoustic model training

2020

Automatic Speech Recognition (ASR) requires huge amounts of real user speech data to reach state-of-the-art performance. However, speech data conveys sensitive speaker attributes like identity that can be inferred and exploited for malicious purposes. Therefore, there is an interest in the collection of anonymized speech data that is processed by some voice conversion method. In this paper, we evaluate one of the voice conversion methods on Latvian speech data and also investigate if privacy-transformed data can be used to improve ASR acoustic models. Results show the effectiveness of voice conversion against state-of-the-art speaker verification models on Latvian speech and the effectivene…

Speaker verificationevaluationvoice conversionComputer scienceSpeech recognitionautomatic speech recognitionLatvianAcoustic model[INFO.INFO-LG] Computer Science [cs]/Machine Learning [cs.LG]privacylanguage.human_language[INFO.INFO-CL]Computer Science [cs]/Computation and Language [cs.CL]anonymization[INFO.INFO-LG]Computer Science [cs]/Machine Learning [cs.LG][INFO.INFO-CL] Computer Science [cs]/Computation and Language [cs.CL]Identity (object-oriented programming)languageConversion methodautomatic speaker verification
researchProduct

Speech Activity Detection under Adverse Noisy Conditions at Low SNRs

2021

Speech originating from the noisy environments degrades the speech quality and intelligibility, thus reducing the human perceived Quality of Experience (QoE). For example, surveillance using drone during natural catastrophe needs an efficient speech recognition device to recognise the speech of the frozen human in presence of drone noise to save their life. Therefore, it often requires to pre-process the noisy speech in order to reduce the noise artifacts and enhance the speech. This paper detects the speech activity using Voice Activity Detection (VAD). The VAD distinguishes speech activity (speech presence) and speech inactivity (silence/noise) by extracting the speech features and compar…

Speech enhancementEuclidean distanceNoiseVoice activity detectionNoise measurementComputer scienceSpeech recognitionFeature extractionSpectral centroidIntelligibility (communication)2021 6th International Conference on Communication and Electronics Systems (ICCES)
researchProduct

Lexical and sublexical units in speech perception.

2009

Saffran, Newport, and Aslin (1996a) found that human infants are sensitive to statistical regularities corresponding to lexical units when hearing an artificial spoken language. Two sorts of segmentation strategies have been proposed to account for this early word-segmentation ability: bracketing strategies, in which infants are assumed to insert boundaries into continuous speech, and clustering strategies, in which infants are assumed to group certain speech sequences together into units (Swingley, 2005). In the present study, we test the predictions of two computational models instantiating each of these strategies i.e., Serial Recurrent Networks: Elman, 1990; and Parser: Perruchet & Vint…

Speech perceptionParsingbusiness.industryCognitive NeuroscienceSpeech recognitionText segmentationExperimental and Cognitive Psychologycomputer.software_genreLexiconSpeech segmentationArtificial Intelligence[SCCO.PSYC]Cognitive science/PsychologyLexicoArtificial intelligenceCluster analysisPsychologybusinesscomputerNatural language processingComputingMilieux_MISCELLANEOUScomputer.programming_languageSpoken languageCognitive science
researchProduct

Interaction in Spoken Word Recognition Models: Feedback Helps

2018

Human perception, cognition, and action requires fast integration of bottom-up signals with top-down knowledge and context. A key theoretical perspective in cognitive science is the interactive activation hypothesis: forward and backward flow in bidirectionally connected neural networks allows humans and other biological systems to approximate optimal integration of bottom-up and top-down information under real-world constraints. An alternative view is that online feedback is neither necessary nor helpful; purely feed forward alternatives can be constructed for any feedback system, and online feedback could not improve processing and would preclude veridical perception. In the domain of spo…

Speech perceptionmedia_common.quotation_subjectSpeech recognitionlcsh:BF1-990Context (language use)speech perception050105 experimental psychologyPsycholinguistics03 medical and health sciences0302 clinical medicinePerceptionspoken word recognition0501 psychology and cognitive sciencesGeneral PsychologypsycholinguisticsBayesian modelsmedia_commonTRACE (psycholinguistics)Computational modelArtificial neural network05 social sciencesFeed forwardlcsh:PsychologySspoken word recognitioncomputational modelssimulationsPsychology030217 neurology & neurosurgeryFrontiers in Psychology
researchProduct

Reading traffic signs while driving: Are linguistic word properties relevant in a complex, dynamic environment?

2019

When driving a vehicle, do we read the words displayed on traffic signs just as we do in more standard conditions? In the driving context, stimulus quality is generally worse, and reading has to be performed at the same time as we are doing other tasks. In the present work, we examined the effects of word frequency and word length on reading in such circumstances. A stimulus presentation mimicking the approach to the traffic sign increased the effect of word frequency, but not the effect of word length, on reading latency. In addition, performing the reading task while driving along a simulated route produced similar results. Therefore, in the context of the driving activity, the advantage …

Speech recognition05 social sciences050109 social psychologyExperimental and Cognitive PsychologyVariable-message signStimulus (physiology)050105 experimental psychologyClinical PsychologyWord lists by frequency0501 psychology and cognitive sciencesPsychologyTraffic signWord lengthApplied PsychologyJournal of Applied Research in Memory and Cognition
researchProduct

Author response: Individual differences in selective attention predict speech identification at a cocktail party

2016

Speech recognitionCocktail partySpeech identificationSelective attentionPsychology
researchProduct

Spatiotemporal Dynamics of the Processing of Spoken Inflected and Derived Words:A Combined EEG and MEG Study

2011

The spatiotemporal dynamics of the neural processing of spoken morphologically complex words are still an open issue. In the current study, we investigated the time course and neural sources of spoken inflected and derived words using simultaneously recorded electroencephalography (EEG) and magnetoencephalography (MEG) responses. Ten participants (native speakers) listened to inflected, derived, and monomorphemic Finnish words and judged their acceptability. EEG and MEG responses were time-locked to both the stimulus onset and the critical point (suffix onset for complex words, uniqueness point for monomorphemic words). The ERP results showed that inflected words elicited a larger left-late…

Speech recognitionElectroencephalographyStimulus (physiology)Lexiconcomputer.software_genre050105 experimental psychologylcsh:RC321-57103 medical and health sciencesBehavioral Neuroscience0302 clinical medicineMorphememorphologymedicine0501 psychology and cognitive sciencesauditorylcsh:Neurosciences. Biological psychiatry. NeuropsychiatryBiological PsychiatryOriginal ResearchTemporal cortexMEGmedicine.diagnostic_testbusiness.industry05 social sciencesderivedMagnetoencephalographyPsychiatry and Mental healthNeuropsychology and Physiological PsychologyNeurologyTime courselexiconArtificial intelligenceSuffixinfectedbusinessPsychologycomputer030217 neurology & neurosurgeryNatural language processingERPNeuroscience
researchProduct