Search results for "Speech recognition"

showing 10 items of 357 documents

A speech recognition approach for an industrial training station

2021

This paper presents a speech recognition service used in the context of commanding and guiding the activities around an industrial training station. The entire concept is built on a decentralized microservice architecture and one of the many hardware and software components is the speech recognition engine. This engine grants users the possibility to interact seamlessly with other components in order to ensure a gradual and productive learning process. By working with different API’s for both English and Romanian languages, the presented approach manages to obtain good speech recognition for defining task phrases aiding the training procedure and to reduce the recognition required time by a…

Service (systems architecture)Process (engineering)Order (business)Speech recognitionComponent-based software engineeringContext (language use)TA1-2040ArchitectureEngineering (General). Civil engineering (General)Training (civil)Task (project management)MATEC Web of Conferences
researchProduct

Modeling musical attributes to characterize ensemble recordings using rhythmic audio features

2011

In this paper, we present the results of a pre-study on music performance analysis of ensemble music. Our aim is to implement a music classification system for the description of live recordings, for instance to help musicologist and musicians to analyze improvised ensemble performances. The main problem we deal with is the extraction of a suitable set of audio features from the recorded instrument tracks. Our approach is to extract rhythm-related audio features and to apply them for regression-based modeling of eight more general musical attributes. The model based on Partial Least-Squares Regression without preceding Principal Component Analysis performed best for all of the eight attribu…

Set (abstract data type)Sound recording and reproductionMusicologyComputer scienceSpeech recognitionFeature extractionMusicalAudio signal processingcomputer.software_genrecomputer2011 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
researchProduct

A Medium Level Language for Pyramid Architectures

1989

In the paper a Parallel C Languages for pyramid architectures is described. The concept of context is introduced in order to handle concurrence between processes in massive parallel machines. Feature implementation on the PAPIA-machine are given.

Settore INF/01 - InformaticaComputer scienceSpeech recognitionConcurrencyPyramidFeature (machine learning)ConcurrenceContext (language use)Parallel computingParallel languages Concurrency Image Analysis Pyramids.
researchProduct

Efficient FPGA Implementation of a Knowledge-Based Automatic Speech Classifier

2005

Speech recognition has become common in many application domains, from dictation systems for professional practices to vocal user interfaces for people with disabilities or hands-free system control. However, so far the performance of Automatic Speech Recognition (ASR) systems are comparable to Human Speech Recognition (HSR) only under very strict working conditions, and in general far lower. Incorporating acoustic-phonetic knowledge into ASR design has been proven a viable approach to rise ASR accuracy. Manner of articulation attributes such as vowel, stop, fricative, approximant, nasal, and silence are examples of such knowledge. Neural networks have already been used successfully as dete…

Settore ING-INF/05 - Sistemi Di Elaborazione Delle InformazioniArtificial neural networkDictationComputer sciencebusiness.industrySpeech recognitionField programmable gate arrays (FPGA)artificial neuralPerceptronManner of articulationKnowledge baseUser interfacebusinessField-programmable gate arrayClassifier (UML)Neural networks
researchProduct

Application of EαNets to Feature Recognition of Articulation Manner in Knowledge-Based Automatic Speech Recognition

2006

Speech recognition has become common in many application domains. Incorporating acoustic-phonetic knowledge into Automatic Speech Recognition (ASR) systems design has been proven a viable approach to rise ASR accuracy. Manner of articulation attributes such as vowel, stop, fricative, approximant, nasal, and silence are examples of such knowledge. Neural networks have already been used successfully as detectors for manner of articulation attributes starting from representations of speech signal frames. In this paper, a set of six detectors for the above mentioned attributes is designed based on the E-αNet model of neural networks. This model was chosen for its capability to learn hidden acti…

Settore ING-INF/05 - Sistemi Di Elaborazione Delle InformazioniArtificial neural networkGeneralizationComputer scienceSpeech recognitionSIGNAL (programming language)cognitive architectureFeature recognitionneural networks speech recognitionAnthropomorphic robotsManner of articulationSystems designSet (psychology)Articulation (phonetics)Robots
researchProduct

Exploiting Correlation between Body Gestures and Spoken Sentences for Real-time Emotion Recognition

2017

Humans communicate their affective states through different media, both verbal and non-verbal, often used at the same time. The knowledge of the emotional state plays a key role to provide personalized and context-related information and services. This is the main reason why several algorithms have been proposed in the last few years for the automatic emotion recognition. In this work we exploit the correlation between one's affective state and the simultaneous body expressions in terms of speech and gestures. Here we propose a system for real-time emotion recognition from gestures. In a first step, the system builds a trusted dataset of association pairs (motion data -> emotion pattern), a…

Settore ING-INF/05 - Sistemi Di Elaborazione Delle InformazioniGround truthSettore INF/01 - InformaticaExploitK-nearest neighborbusiness.industrySpeech recognitioncomputer.software_genreMotion (physics)CorrelationDynamic Time Warping Emotion Recognition K-nearest neighborEmotion RecognitionKey (cryptography)Artificial intelligenceState (computer science)businessAssociation (psychology)PsychologycomputerNatural language processingGestureDynamic Time Warping
researchProduct

InspirationWall

2015

Collaborative idea generation leverages social interactions and knowledge sharing to spark diverse associations and produce creative ideas. Information exploration systems expand the current context by suggesting novel but related concepts. In this paper we introduce InspirationWall, an unobtrusive display that leverages speech recognition and information exploration to enhance an ongoing idea generation session with automatically retrieved concepts that relate to the conversation. We evaluated the system in six idea generation sessions of 20 minutes with small groups of two people. Preliminary results suggest that InspirationWall contrasts the decay of idea productivity over time and can t…

Settore ING-INF/05 - Sistemi Di Elaborazione Delle InformazioniInformation ExplorationSettore INF/01 - InformaticaComputer sciencemedia_common.quotation_subjectContext (language use)Automatic Speech RecognitionIdeationIdea generationSession (web analytics)Knowledge sharingSPARK (programming language)Human–computer interactionConversationInformation explorationcomputercomputer.programming_languagemedia_commonProceedings of the 2015 ACM SIGCHI Conference on Creativity and Cognition
researchProduct

Keyword Based Keyframe Extraction in Online Video Collections

2015

Keyframe extraction methods aim to find in a video sequence the most significant frames, according to specific criteria. In this paper we propose a new method to search, in a video database, for frames that are related to a given keyword, and to extract the best ones, according to a proposed quality factor. We first exploit a speech to text algorithm to extract automatic captions from all the video in a specific domain database. Then we select only those sequences (clips), whose captions include a given keyword, thus discarding a lot of information that is useless for our purposes. Each retrieved clip is then divided into shots, using a video segmentation method, that is based on the SURF d…

Settore ING-INF/05 - Sistemi Di Elaborazione Delle InformazioniInformation retrievalbusiness.industryComputer sciencemedia_common.quotation_subjectShot (filmmaking)InformationSystems_INFORMATIONSTORAGEANDRETRIEVALFrame (networking)ComputingMethodologies_IMAGEPROCESSINGANDCOMPUTERVISIONPattern recognitionDomain (software engineering)Factor (programming language)Metric (mathematics)Quality (business)SegmentationArtificial intelligencebusinesscomputerSentencemedia_commoncomputer.programming_languageVideo Summarization Keyframe Extraction Automatic Speech Recognition YouTube Multimedia Collections
researchProduct

Real-Time Body Gestures Recognition Using Training Set Constrained Reduction

2017

Gesture recognition is an emerging cross-discipline research field, which aims at interpreting human gestures and associating them to a well-defined meaning. It has been used as a mean for supporting human to machine interaction in several applications of robotics, artificial intelligence, and machine learning. In this paper, we propose a system able to recognize human body gestures which implements a constrained training set reduction technique. This allows the system for a real-time execution. The system has been tested on a publicly available dataset of 7,000 gestures, and experimental results have highlighted that at the cost of a little decrease in the maximum achievable recognition ac…

Settore ING-INF/05 - Sistemi Di Elaborazione Delle InformazioniTraining setComputer sciencebusiness.industrySpeech recognitionRoboticsGesture Recognition Real-time systems Constrained optimizationField (computer science)Reduction (complexity)Gesture recognitionArtificial intelligencebusinessGestureMeaning (linguistics)
researchProduct

Embedded Knowledge-based Speech Detectors for Real-Time Recognition Tasks

2006

Speech recognition has become common in many application domains, from dictation systems for professional practices to vocal user interfaces for people with disabilities or hands-free system control. However, so far the performance of automatic speech recognition (ASR) systems are comparable to human speech recognition (HSR) only under very strict working conditions, and in general much lower. Incorporating acoustic-phonetic knowledge into ASR design has been proven a viable approach to raise ASR accuracy. Manner of articulation attributes such as vowel, stop, fricative, approximant, nasal, and silence are examples of such knowledge. Neural networks have already been used successfully as de…

Settore ING-INF/05 - Sistemi Di Elaborazione Delle InformazioniVoice activity detectionArtificial neural networkDictationbusiness.industryComputer scienceSpeech recognitionSpeech technologycomputer.software_genreSpeech processingManner of articulationSilenceVowelComputer ScienceTelecommunicationsMel-frequency cepstrumArtificial intelligencespeech detectorUser interfacebusinesscomputerNatural language processing
researchProduct