Empresas
Empregos
  • Sobre nós
  • Soluções
    • Publicação de vagas
      Publique sua vaga e receba candidatos qualificados em 48h.
    • Avaliações de candidatos
      Mais de 500 testes técnicos e psicológicos, mais anti-fraude.
    • Headhunting
      Busca executiva personalizada do início ao fim.
    • Folha de Pagamento + EOR
      Dispersão de folha e EOR em mais de 15 países da LATAM.
  • Preços
  • Empregos

0

230
Visualizações
Different S3 behavior using different endpoints?

I'm currently writing code to use Amazon's S3 REST API and I notice different behavior where the only difference seems to be the Amazon endpoint URI that I use, e.g., https://s3.amazonaws.com vs. https://s3-us-west-2.amazonaws.com.

Examples of different behavior for the the GET Bucket (List Objects) call:

  • Using one endpoint, it includes the "folder" in the results, e.g.:

    /path/subfolder/
    /path/subfolder/file1.txt
    /path/subfolder/file2.txt
    

    and, using the other endpoint, it does not include the "folder" in the results:

    /path/subfolder/file1.txt
    /path/subfolder/file2.txt
    
  • Using one endpoint, it represents "folders" using a trailing / as shown above and, using the other endpoint, it uses a trailing _$folder$:

    /path/subfolder_$folder$
    /path/subfolder/file1.txt
    /path/subfolder/file2.txt
    

Why the differences? How can I make it return results in a consistent manner regardless of endpoint?

Note that I get these same odd results even if I use Amazon's own command-line AWS S3 client, so it's not my code.

about 4 years ago · Santiago Trujillo
3 Respostas
Responde à pergunta

0

And the contents of the buckets should be irrelevant anyway.

Your assertion notwithstanding, your issue is exactly about the content of the buckets, and not something S3 is doing -- the S3 API has no concept of folders. None. The S3 console can display folders, but this is for convenience -- the folders are not really there -- or if there are folder-like entities, they're irrelevant and not needed.

In Amazon S3, buckets and objects are the primary resources, where objects are stored in buckets. Amazon S3 has a flat structure with no hierarchy like you would see in a typical file system. However, for the sake of organizational simplicity, the Amazon S3 console supports the folder concept as a means of grouping objects. Amazon S3 does this by using key name prefixes for objects.

http://docs.aws.amazon.com/AmazonS3/latest/UG/FolderOperations.html

So why are you seeing this?

Either you've been using EMR/Hadoop, or some other code written by someone who took a bad example and ran with it... or is doing something differently than it should have been done for quite some time.

Amazon EMR is a web service that uses a managed Hadoop framework to process, distribute, and interact with data in AWS data stores, including Amazon S3. Because S3 uses a key-value pair storage system, the Hadoop file system implements directory support in S3 by creating empty files with the <directoryname>_$folder$ suffix.

https://aws.amazon.com/premiumsupport/knowledge-center/emr-s3-empty-files/

This may have been something the S3 console did many years ago, and apparently (since you don't report seeing them in the console) it still supports displaying such objects as folders in the console... but the S3 console no longer creates them this way, if it ever did.

I've mirrored the bucket "folder" layout exactly

If you create a folder in the console, an empty object with the key "foldername/" is created. This in turn is used to display a folder that you can navigate into, and upload objects with keys beginning with that folder name as a prefix.

The Amazon S3 console treats all objects that have a forward slash "/" character as the last (trailing) character in the key name as a folder

http://docs.aws.amazon.com/AmazonS3/latest/UG/FolderOperations.html

If you just create objects using the API, then "my/object.txt" appears in the console as "object.txt" inside folder "my" even though there is no "my/" object created... so if the objects are created with the API, you'd see neither style of "folder" in the object listing.

about 4 years ago · Santiago Trujillo Relatório

0

That is probably a bug in the API endpoint which includes the "folder" - S3 internally doesn't actually have a folder structure, but instead is just a set of keys associated with files, where keys (for convenience) can contain slash-separated paths which then show up as "folders" in the web interface. There is the option in the API to specify a prefix, which I believe can be any part of the key up to and including part of the filename.

about 4 years ago · Santiago Trujillo Relatório

0

EMR's s3 client is not the apache one, so I can't speak accurately about it.

In ASF hadoop releases (and HDP, CDH)

  1. The older s3n:// client uses $folder$ as its folder delimiter.
  2. The newer s3a:// client uses / as its folder marker, but will handle $folder$ if there. At least it used to; I can't see where in the code it does now.

The S3A clients strip out all folder markers when you list things; S3A uses them to simulate empty dirs and deletes all parent markers when you create child file/dir entries.

Whatever you have which processes GET should just ignore entries with "/" or $folder at the end.

As to why they are different, the local EMRFS is a different codepath, using dynamo for implementing consistency. At a guess, it doesn't need to mock empty dirs, as the DDB tables will host all directory entries.

about 4 years ago · Santiago Trujillo Relatório
Responde à pergunta
Encontrar trabalhos remotos

Descubra a nova forma de encontrar um emprego!

melhores empregos
Principais categorias de trabalho
Empresas
Postar vaga Preços Comercial
Jurídico
Termos e Condições Política de privacidade
© 2026 PeakU Inc. All Rights Reserved.
Andres GPT
Recomende algumas ofertas para mim
Preciso de ajuda