Business
Jobs
  • About Us
  • Solutions
    • Job Postings
      Post your job and receive qualified candidates in 48h.
    • Candidate Assessments
      500+ technical and psychological tests, plus anti-fraud.
    • Headhunting
      Tailor-made executive search from start to finish.
    • Payroll + EOR
      Payroll dispersal and EOR across 15+ LATAM countries.
  • Pricing
  • Jobs

0

142
Views
¿Alguien conoce una buena heurística para detectar UTF-8 mal decodificado como texto Latin-1?

Recibo alertas meteorológicas de un servicio meteorológico. Aunque la respuesta HTTP dice ser UTF-8, claramente contiene texto como este:

Suðaustan 13-20 m/s og snjókoma með lélegu skyggni og versnandi akstursskilyrðum.

... que debería verse así:

Suðaustan 13-20 m/s og snjókoma með lélegu skyggni og versnandi akstursskilyrðum.

... pero ya se decodificó incorrectamente antes de que me llegara por primera vez, se volvió a codificar como UTF-8 después de decodificarse incorrectamente. La mayoría de nosotros probablemente haya visto este tipo de basura "mojibake" antes, y al menos visualmente, a menudo tiene muchas características comunes, como muchos caracteres à , signos ¢ y similares.

Estoy usando este código para arreglarlo ahora mismo:

 // Check for UTF-8 wrongly decoded as Latin-1 if (/[\x80-\xC5]/.test(result)) { const bytes = Buffer.from(result, 'latin1'); const altText = bytes.toString('utf8'); if (altText.length < result.length) result = altText; }

...y eso está funcionando por ahora, pero no es una prueba muy sofisticada.

¿Alguien sabe de un método mejor?

about 4 years ago · Juan Pablo Isaza
1 answers
Answer question

0

¿Alguien sabe de un método mejor?

No sé cómo determinaría mejor. Escribí esta función hace un tiempo para hacer exactamente esta transformación en una cadena.

No sé si eso es mejor que el Buffer.

 function utf8_decode(str) { //assuming the input is a valid utf-8 string. //Invalid parts are ignored / remain in the string. return str.replace( /[\u00c0-\u00df][\u0080-\u00bf]|([\u00e0-\u00ef][\u0080-\u00bf]{2})|([\u00f0-\u00f7][\u0080-\u00bf]{3})/g, (two, three, four) => String.fromCodePoint( // UTF-16 codePoints four ? (four.charCodeAt(0) & 7) << 18 | (four.charCodeAt(1) & 63) << 12 | (four.charCodeAt(2) & 63) << 6 | (four.charCodeAt(3) & 63) : // UTF-8 multibytes three ? (three.charCodeAt(0) & 15) << 12 | (three.charCodeAt(1) & 63) << 6 | (three.charCodeAt(2) & 63) : (two.charCodeAt(0) & 31) << 6 | (two.charCodeAt(1) & 63) ) ) } console.log(utf8_decode("Suðaustan 13-20 m/s og snjókoma með lélegu skyggni og versnandi akstursskilyrðum.")); console.log(utf8_decode("ð\x9F\x98\x8B"));

La expresión regular es mejor que la tuya.

No es necesario verificar después si la transformación resultó en algún cambio.

about 4 years ago · Juan Pablo Isaza Report
Answer question
Find remote jobs

Discover the new way to find a job!

Top jobs
Top job categories
Business
Post vacancy Pricing Sales
Legal
Terms and conditions Privacy policy
© 2026 PeakU Inc. All Rights Reserved.
Andres GPT
Show me some job opportunities
There's an error!