Business
Jobs
  • About Us
  • Solutions
    • Job Postings
      Post your job and receive qualified candidates in 48h.
    • Candidate Assessments
      500+ technical and psychological tests, plus anti-fraud.
    • Headhunting
      Tailor-made executive search from start to finish.
    • Payroll + EOR
      Payroll dispersal and EOR across 15+ LATAM countries.
  • Pricing
  • Jobs

0

538
Views
PySpark reemplaza nulo en la columna con valor en otra columna

Quiero reemplazar los valores nulos en una columna con los valores en una columna adyacente, por ejemplo, si tengo

 A|B 0,1 2,null 3,null 4,2

quiero que sea:

 A|B 0,1 2,2 3,3 4,2

probado con

 df.na.fill(df.A,"B")

Pero no funcionó, dice que el valor debe ser float, int, long, string o dict

¿Algunas ideas?

over 4 years ago · Santiago Trujillo
3 answers
Answer question

0

Podemos usar coalesce

 from pyspark.sql.functions import coalesce df.withColumn("B",coalesce(df.B,df.A))
over 4 years ago · Santiago Trujillo Report

0

Otra respuesta.

Si el siguiente df1 su marco de datos

 rd1 = sc.parallelize([(0,1), (2,None), (3,None), (4,2)]) df1 = rd1.toDF(['A', 'B']) from pyspark.sql.functions import when df1.select('A', when( df1.B.isNull(), df1.A).otherwise(df1.B).alias('B') )\ .show()
over 4 years ago · Santiago Trujillo Report

0

df.rdd.map(lambda row: row if row[1] else Row(a=row[0],b=row[0])).toDF().show()
over 4 years ago · Santiago Trujillo Report
Answer question
Find remote jobs

Discover the new way to find a job!

Top jobs
Top job categories
Business
Post vacancy Pricing Sales
Legal
Terms and conditions Privacy policy
© 2026 PeakU Inc. All Rights Reserved.
Andres GPT
Show me some job opportunities
There's an error!