Empresas
Empregos
  • Sobre nós
  • Soluções
    • Publicação de vagas
      Publique sua vaga e receba candidatos qualificados em 48h.
    • Avaliações de candidatos
      Mais de 500 testes técnicos e psicológicos, mais anti-fraude.
    • Headhunting
      Busca executiva personalizada do início ao fim.
    • Folha de Pagamento + EOR
      Dispersão de folha e EOR em mais de 15 países da LATAM.
  • Preços
  • Empregos

0

237
Visualizações
Why are my ORC with SNAPPY compressed files larger than the original files?

I set up a first Hive table with GZIP-compressed files:

CREATE EXTERNAL TABLE table_gzip (
    col1,
    col2,
    col3
)
ROW FORMAT DELIMITED,
  FIELDS TERMINATED BY ','
  LINES TERMINATED BY '\n'
LOCATION
  's3://bucket/files_gzip/';

Then I set up another Hive table with ORC format:

CREATE EXTERNAL TABLE table_orc (
    col1,
    col2,
    col3
)
STORED AS ORC
LOCATION
   's3://bucket/files_orc/';
ALTER TABLE table_orc SET tblproperties ("orc.compress" ="SNAPPY");

And then I uncompressed and recompressed from GZIP to ORC using this query:

INSERT OVERWRITE TABLE table_gzip SELECT * FROM table_orc

Once this query finished, I had new ORC-compressed files in 's3://bucket/files_orc/'. So far so good.

However when I looked at the files, they went from 500 1.2GiB files to 500 1.6GiB files.

What did I do wrong? Why are my ORC-SNAPPY compressed files larger than the original files? Is GZIP a better compression method?

Thanks for your time.

over 4 years ago · Santiago Trujillo
Responde à pergunta
Encontrar trabalhos remotos

Descubra a nova forma de encontrar um emprego!

melhores empregos
Principais categorias de trabalho
Empresas
Postar vaga Preços Comercial
Jurídico
Termos e Condições Política de privacidade
© 2026 PeakU Inc. All Rights Reserved.
Andres GPT
Recomende algumas ofertas para mim
Preciso de ajuda