Empresas
Empleos
  • Sobre nosotros
  • Soluciones
    • Publicación de vacantes
      Publica tu vacante y recibe candidatos calificados en 48h.
    • Evaluación de candidatos
      500+ pruebas técnicas y psicológicas, más anti-fraude.
    • Headhunting
      Búsqueda ejecutiva a la medida de principio a fin.
    • Nómina + EOR
      Dispersión de nómina y EOR en más de 15 países de LATAM.
  • Precios
  • Empleos

0

217
Vistas
Save speach used to fill a textbox after dictating on (ipad/mobile device) with keyboard-microphone

By clicking on a textbox on ipad or mobile devices in a browser the keyboard appears on the screen. Here is possible to select the microphone to dictate the text directly into the inputbox with our voice without the need to write directly.

Because the speach conversion is not always perfet we vould like to save the audio of the speach itself on our server to be used when the text is not clear enough.

Is it possible to retrive, from the ipad/mobile, and save on our server, the audio of the speach that has been used to write the text in out textbox?

I know that i could write javascript code to convert the speach in text and record the voice directly but we would like to know if it is possible to get the audio (as a file) used for the speachtotext conversion used by the device to fill textbox.

In other words when I dictate by using the microphone, of the device keybord, does the device allows the page, where the coversion took place, to access the audio as a file?

about 4 years ago · Juan Pablo Isaza
1 Respuestas
Responde la pregunta

0

Since you tagged this Android as well- not on that OS. The keyboard is its own app and handles the voice input itself. There is no way to access the files of another app.

If you want to do this in a native app, put up your own microphone button and use the speech to text service, which will return an array of possible inputs with probabilities. In a browser, you're just out of luck as there is no access to the service.

All of this is kind of a moot point anyway for a few reasons

  • Very few people use speech input. My last data on numbers is old, but it was unpopular enough when I worked at a keyboard company we had an option to remove the key.
  • Uploading those files would be a huge privacy concern. Look at the firestorm a year or so ago when it was found Google/Amazon did this same thing for the same reason. This was a bigger deal in their case as it was background processing, but users would likely still not be happy.
  • Unless you're spending a few million on researchers, you're not going to be doing it better than the existing solution. That kind of software is not easy to write, its not a totally solved problem even by Google and Nuance (owners of Dragon which powers Siri, or at least did) who have huge teams. Why do you think you'll do better? Unless you plan on listening to them manually. In which case the next point is even bigger.
  • Ok, so you upload the file and you do find a better solution- what are you going to do about it? Somehow change the text the user typed in 20 minutes ago? How are you going to do this and have a UX flow that makes sense?
about 4 years ago · Juan Pablo Isaza Denunciar
Responde la pregunta
Encuentra empleos remotos

¡Descubre la nueva forma de encontrar empleo!

Top de empleos
Top categorías de empleo
Empresas
Publicar vacante Precios Comercial
Legal
Términos y condiciones Política de privacidad
© 2026 PeakU Inc. All Rights Reserved.
Andres GPT
Recomiéndame algunas ofertas
Necesito ayuda