Business
Jobs
  • About Us
  • Solutions
    • Job Postings
      Post your job and receive qualified candidates in 48h.
    • Candidate Assessments
      500+ technical and psychological tests, plus anti-fraud.
    • Headhunting
      Tailor-made executive search from start to finish.
    • Payroll + EOR
      Payroll dispersal and EOR across 15+ LATAM countries.
  • Pricing
  • Jobs

0

132
Views
Lea csv de Amazon s3 usando python2.7

Puedo obtener fácilmente el nombre del depósito de s3, pero cuando leo el archivo csv de s3, siempre aparece un error.

 import boto3 import pandas as pd s3 = boto3.client('s3', aws_access_key_id='yyyyyyyy', aws_secret_access_key='xxxxxxxxxxx') # Call S3 to list current buckets response = s3.list_buckets() for bucket in response['Buckets']: print bucket['Name'] output s3-bucket-data

.

 import pandas as pd import StringIO from boto.s3.connection import S3Connection AWS_KEY = 'yyyyyyyyyy' AWS_SECRET = 'xxxxxxxxxx' aws_connection = S3Connection(AWS_KEY, AWS_SECRET) bucket = aws_connection.get_bucket('s3-bucket-data') fileName = "data.csv" content = bucket.get_key(fileName).get_contents_as_string() reader = pd.read_csv(StringIO.StringIO(content))

recibiendo error-

 boto.exception.S3ResponseError: S3ResponseError: 400 Bad Request

¿Cómo puedo leer el csv desde s3?

about 4 years ago · Santiago Trujillo
3 answers
Answer question

0

puedes usar el paquete s3fs

s3fs también admite perfiles aws en archivos de credenciales.

Aquí hay un ejemplo (no es necesario que lo fragmente, pero solo tenía este ejemplo a mano),

 import os import pandas as pd import s3fs import gzip chunksize = 999999 usecols = ["Col1", "Col2"] filename = 'some_csv_file.csv.gz' s3_bucket_name = 'some_bucket_name' AWS_KEY = 'yyyyyyyyyy' AWS_SECRET = 'xxxxxxxxxx' s3f = s3fs.S3FileSystem( anon=False, key=AWS_KEY, secret=AWS_SECRET) # or if you have a profile defined in credentials file: #aws_shared_credentials_file = 'path/to/aws/credentials/file/' #os.environ['AWS_SHARED_CREDENTIALS_FILE'] = aws_shared_credentials_file #s3f = s3fs.S3FileSystem( # anon=False, # profile_name=s3_profile) filepath = os.path.join(s3_bucket_name, filename) with s3f.open(filepath, 'rb') as f: gz = gzip.GzipFile(fileobj=f) # Decompress data with gzip chunks = pd.read_csv(gz, usecols=usecols, chunksize=chunksize, iterator=True, ) df = pd.concat([c for c in chunks], axis=1)
about 4 years ago · Santiago Trujillo Report

0

boto es algo que me encanta cuando se trata de manejar datos en S3 con python.

instalar boto usando pip install boto

 import boto from boto.s3.key import Key keyId ="your_aws_key_id" sKeyId="your_aws_secret_key_id" srcFileName="abc.txt" # filename on S3 destFileName="s3_abc.txt" # output file name bucketName="mybucket001" # S3 bucket name conn = boto.connect_s3(keyId,sKeyId) bucket = conn.get_bucket(bucketName) #Get the Key object of the given key, in the bucket k = Key(bucket,srcFileName) #Get the contents of the key into a file k.get_contents_to_filename(destFileName)
about 4 years ago · Santiago Trujillo Report

0

Experimenté este problema con algunas regiones de AWS. Creé un depósito en "us-east-1" y el siguiente código funcionó bien:

 import boto from boto.s3.key import Key import StringIO import pandas as pd keyId ="xxxxxxxxxxxxxxxxxx" sKeyId="yyyyyyyyyyyyyyyyyy" srcFileName="zzzzz.csv" bucketName="elasticbeanstalk-us-east-1-aaaaaaaaaaaa" conn = boto.connect_s3(keyId,sKeyId) bucket = conn.get_bucket(bucketName) k = Key(bucket,srcFileName) content = k.get_contents_as_string() reader = pd.read_csv(StringIO.StringIO(content))

Intente crear un depósito nuevo en us-east-1 y vea si funciona.

about 4 years ago · Santiago Trujillo Report
Answer question
Find remote jobs

Discover the new way to find a job!

Top jobs
Top job categories
Business
Post vacancy Pricing Sales
Legal
Terms and conditions Privacy policy
© 2026 PeakU Inc. All Rights Reserved.
Andres GPT
Show me some job opportunities
There's an error!