Search results for "Speech recognition"

showing 7 items of 357 documents

miMic

2016

miMic, a sonic analogue of paper and pencil is proposed: An augmented microphone for vocal and gestural sonic sketching. Vocalizations are classified and interpreted as instances of sound models, which the user can play with by vocal and gestural control. The physical device is based on a modified microphone, with embedded inertial sensors and buttons. Sound models can be selected by vocal imitations that are automatically classified, and each model is mapped to vocal and gestural features for real-time control. With miMic, the sound designer can explore a vast sonic space and quickly produce expressive sonic sketches, which may be turned into sound prototypes by further adjustment of model…

sonic interaction design system architecture vocal sketching sound designSound (medical instrument)Settore INF/01 - InformaticaInformationSystems_INFORMATIONINTERFACESANDPRESENTATION(e.g.HCI)Computer scienceMicrophoneSound designSpeech recognition05 social sciencesModel parameterssystem architecturevocal sketchingsonic interaction designAugmented microphone050105 experimental psychologysound designGestureInertial measurement unitSonic interaction designSettore ICAR/13 - Disegno Industriale0501 psychology and cognitive sciences050107 human factorsPencil (mathematics)GestureProceedings of the TEI '16: Tenth International Conference on Tangible, Embedded, and Embodied Interaction
researchProduct

Soundscape design through evolutionary engines

2008

Abstract Two implementations of an Evolutionary Sound Synthesis method using the Interaural Time Difference (ITD) and psychoacoustic descriptors are presented here as a way to develop criteria for fitness evaluation. We also explore a relationship between adaptive sound evolution and three soundscape characteristics: keysounds, key-signals and sound-marks. Sonic Localization Field is defined using a sound attenuation factor and ITD azimuth angle, respectively (Ii, Li). These pairs are used to build Spatial Sound Genotypes (SSG) and they are extracted from a waveform population set. An explanation on how our model was initially written in MATLAB is followed by a recent Pure Data (Pd) impleme…

sonic spatializationeducation.field_of_studySoundscapesound synthesisGeneral Computer Scienceartificial evolutionComputer scienceSpeech recognitionacoustic descriptorsPopulationEvolutionary algorithmInteraural time differencegenetic algorithmsPure DataPsychoacousticseducationcomputerAcoustic attenuationParametric statisticscomputer.programming_languageComputer Science(all)
researchProduct

Optimal Volume for Concert Halls Based on Ando’s Subjective Preference and Barron Revised Theories

2014

[EN] The Ando-Beranek s model, a linear version of Ando s subjective preference theory, obtained by the authors in a recent work, was combined with Barron revised theory. An optimal volume region for each reverberation time was obtained for classical music in symphony orchestra concert halls. The obtained relation was tested with good agreement with the top rated halls reported by Beranek and other halls with reported anomalies.

sound qualitySound qualityoptimal criteriumSpeech recognitionOptimal criteriumBuilding and ConstructionRoom acousticslcsh:TH1-9745Classical musicPreference theoryFISICA APLICADAroom acousticsArchitectureSymphonyMATEMATICA APLICADAPsychologyRoom acousticsPreference (economics)Mathematical economicsroom acoustics; sound quality; optimal criteriumlcsh:Building constructionCivil and Structural EngineeringBuildings
researchProduct

A multimodal chat-bot based information technology system

2006

The proposed system integrates chat-bot and speech recognition technologies in order to build a versatile, user-friendly, virtual assistant guide with information retrieval capabilities. The system is adaptable to the user needs of mobility being also usable on different devices (i.e. PDAs, Smartphone). The system has been implemented on a Qtek 9090 with Windows Mobile 2003 and a simulation for the cultural heritage domain is here presented.

speech synthesisman-machine interfacesspeech recognition
researchProduct

Exploiting ongoing EEG with multilinear partial least squares during free-listening to music

2016

During real-world experiences, determining the stimulus-relevant brain activity is excitingly attractive and is very challenging, particularly in electroencephalography. Here, spectrograms of ongoing electroencephalogram (EEG) of one participant constructed a third-order tensor with three factors of time, frequency and space; and the stimulus data consisting of acoustical features derived from the naturalistic and continuous music formulated a matrix with two factors of time and the number of features. Thus, the multilinear partial least squares (PLS) conforming to the canonical polyadic (CP) model was performed on the tensor and the matrix for decomposing the ongoing EEG. Consequently, we …

ta113Multilinear mapmedicine.diagnostic_testBrain activity and meditationSpeech recognition02 engineering and technologyElectroencephalographyta3112Matrix decomposition03 medical and health sciences0302 clinical medicinetensor decompositionFrequency domainPartial least squares regression0202 electrical engineering electronic engineering information engineeringmedicineSpectrogramOngoing EEG020201 artificial intelligence & image processingmusicTime domain030217 neurology & neurosurgerymultilinear partial least squaresMathematics
researchProduct

Combining PCA and multiset CCA for dimension reduction when group ICA is applied to decompose naturalistic fMRI data

2015

An extension of group independent component analysis (GICA) is introduced, where multi-set canonical correlation analysis (MCCA) is combined with principal component analysis (PCA) for three-stage dimension reduction. The method is applied on naturalistic functional MRI (fMRI) images acquired during task-free continuous music listening experiment, and the results are compared with the outcome of the conventional GICA. The extended GICA resulted slightly faster ICA convergence and, more interestingly, extracted more stimulus-related components than its conventional counterpart. Therefore, we think the extension is beneficial enhancement for GICA, especially when applied to challenging fMRI d…

ta113MultisetPCAGroup (mathematics)business.industrydimension reductionSpeech recognitionDimensionality reductionPattern recognitionMusic listeningta3112naturalistic fMRIGroup independent component analysisPrincipal component analysistemporal cocatenationArtificial intelligenceCanonical correlationbusinessmultiset CCAMathematics
researchProduct

Visual Distraction Effects of In-Car Text Entry Methods

2017

Three text entry methods were compared in a driving simulator study with 17 participants. Ninety-seven drivers’ occlusion distance (OD) data mapped on the test routes was used as a baseline to evaluate the methods’ visual distraction potential. Only the voice recognition-based text entry tasks passed the set verification criteria. Handwriting tasks were experienced as the most demanding and the voice recognition tasks as the least demanding. An individual in-car glance length preference was found, but against expectations, drivers’ ODs did not correlate with incar glance lengths or visual short-term memory capacity. The handwriting method was further studied with 24 participants with instru…

ta113visual short-term memorydriver distraction050210 logistics & transportationocclusion distanceVisual Patterns TestComputer scienceSpeech recognition05 social sciencesDriving simulatorvisual demandAffect (psychology)Test (assessment)HandwritingDistraction0502 economics and businesstext entry methods0501 psychology and cognitive sciencesVisual short-term memorySet (psychology)050107 human factorsReliability (statistics)visual occlusionProceedings of the 9th International Conference on Automotive User Interfaces and Interactive Vehicular Applications
researchProduct