Empresas
Empregos
  • Sobre nós
  • Soluções
    • Publicação de vagas
      Publique sua vaga e receba candidatos qualificados em 48h.
    • Avaliações de candidatos
      Mais de 500 testes técnicos e psicológicos, mais anti-fraude.
    • Headhunting
      Busca executiva personalizada do início ao fim.
    • Folha de Pagamento + EOR
      Dispersão de folha e EOR em mais de 15 países da LATAM.
  • Preços
  • Empregos

0

216
Visualizações
Save speach used to fill a textbox after dictating on (ipad/mobile device) with keyboard-microphone

By clicking on a textbox on ipad or mobile devices in a browser the keyboard appears on the screen. Here is possible to select the microphone to dictate the text directly into the inputbox with our voice without the need to write directly.

Because the speach conversion is not always perfet we vould like to save the audio of the speach itself on our server to be used when the text is not clear enough.

Is it possible to retrive, from the ipad/mobile, and save on our server, the audio of the speach that has been used to write the text in out textbox?

I know that i could write javascript code to convert the speach in text and record the voice directly but we would like to know if it is possible to get the audio (as a file) used for the speachtotext conversion used by the device to fill textbox.

In other words when I dictate by using the microphone, of the device keybord, does the device allows the page, where the coversion took place, to access the audio as a file?

about 4 years ago · Juan Pablo Isaza
1 Respostas
Responde à pergunta

0

Since you tagged this Android as well- not on that OS. The keyboard is its own app and handles the voice input itself. There is no way to access the files of another app.

If you want to do this in a native app, put up your own microphone button and use the speech to text service, which will return an array of possible inputs with probabilities. In a browser, you're just out of luck as there is no access to the service.

All of this is kind of a moot point anyway for a few reasons

  • Very few people use speech input. My last data on numbers is old, but it was unpopular enough when I worked at a keyboard company we had an option to remove the key.
  • Uploading those files would be a huge privacy concern. Look at the firestorm a year or so ago when it was found Google/Amazon did this same thing for the same reason. This was a bigger deal in their case as it was background processing, but users would likely still not be happy.
  • Unless you're spending a few million on researchers, you're not going to be doing it better than the existing solution. That kind of software is not easy to write, its not a totally solved problem even by Google and Nuance (owners of Dragon which powers Siri, or at least did) who have huge teams. Why do you think you'll do better? Unless you plan on listening to them manually. In which case the next point is even bigger.
  • Ok, so you upload the file and you do find a better solution- what are you going to do about it? Somehow change the text the user typed in 20 minutes ago? How are you going to do this and have a UX flow that makes sense?
about 4 years ago · Juan Pablo Isaza Relatório
Responde à pergunta
Encontrar trabalhos remotos

Descubra a nova forma de encontrar um emprego!

melhores empregos
Principais categorias de trabalho
Empresas
Postar vaga Preços Comercial
Jurídico
Termos e Condições Política de privacidade
© 2026 PeakU Inc. All Rights Reserved.
Andres GPT
Recomende algumas ofertas para mim
Preciso de ajuda