Empresas
Empleos
  • Sobre nosotros
  • Soluciones
    • Publicación de vacantes
      Publica tu vacante y recibe candidatos calificados en 48h.
    • Evaluación de candidatos
      500+ pruebas técnicas y psicológicas, más anti-fraude.
    • Headhunting
      Búsqueda ejecutiva a la medida de principio a fin.
    • Nómina + EOR
      Dispersión de nómina y EOR en más de 15 países de LATAM.
  • Precios
  • Empleos

0

110
Vistas
Sampling from multiple tables

We are currently in the process of writing quite complex spark jobs that contain multiple joins and a filters across multiple tables.

We'd like to unit test these jobs with actual data but the real data lives in the cloud (S3 buckets) spreads over dozens of tables (orc files), millions of rows each.

The inherent problem is that sampling from multiple tables can cause joins on the samples to produce no results since there is a possibility that certain IDs don't exist in both sampled tables. Is there a way (a heuristic or a tool) to sample data (for example 1000 rows) from each tables in such a way that joins on foreign keys still provide a useful amount of rows?

That way we can use the sampled data offline to unit test the jobs.

about 4 years ago · Santiago Trujillo
Responde la pregunta
Encuentra empleos remotos

¡Descubre la nueva forma de encontrar empleo!

Top de empleos
Top categorías de empleo
Empresas
Publicar vacante Precios Comercial
Legal
Términos y condiciones Política de privacidad
© 2026 PeakU Inc. All Rights Reserved.
Andres GPT
Recomiéndame algunas ofertas
Necesito ayuda