Empresas
Empleos
  • Sobre nosotros
  • Soluciones
    • Publicación de vacantes
      Publica tu vacante y recibe candidatos calificados en 48h.
    • Evaluación de candidatos
      500+ pruebas técnicas y psicológicas, más anti-fraude.
    • Headhunting
      Búsqueda ejecutiva a la medida de principio a fin.
    • Nómina + EOR
      Dispersión de nómina y EOR en más de 15 países de LATAM.
  • Precios
  • Empleos

0

302
Vistas
Update a column in a table having huge data (80mn+ rows) in cassandra

I have a table in Cassandra which has almost 80 million+ records(may be more than that). I have updated the schama which adds a new column in the table. Now I need to update the column values. I wrote a migration script to do that using cassandra-driver. Tried batching, token but the data is so huge that it is taking more than 3 hrs and still not updating the records (process getting terminated after 2-3 hrs.) What is the best way to handle this type of migration ? Is there any other way to achieve this?

Token example

over 4 years ago · Santiago Trujillo
1 Respuestas
Responde la pregunta

0

Usually for such things it's easier to use Spark (although I'm not sure hot it works with Amazon Keyspaces). It's quite hard to do range scan correctly - you need to handle edge cases, etc. (I have an example for Java driver that uses the same algorithm as Spark Cassandra Connector and DSBulk).

You can use Python with Spark and Cassandra Connector to update your data - the complexity of update will depend on your algorithm.

Another approach is put logic into your App - if it receives from Cassandra null for given column, you can return calculated value.

over 4 years ago · Santiago Trujillo Denunciar
Responde la pregunta
Encuentra empleos remotos

¡Descubre la nueva forma de encontrar empleo!

Top de empleos
Top categorías de empleo
Empresas
Publicar vacante Precios Comercial
Legal
Términos y condiciones Política de privacidad
© 2026 PeakU Inc. All Rights Reserved.
Andres GPT
Recomiéndame algunas ofertas
Necesito ayuda