Business
Jobs
  • About Us
  • Solutions
    • Job Postings
      Post your job and receive qualified candidates in 48h.
    • Candidate Assessments
      500+ technical and psychological tests, plus anti-fraud.
    • Headhunting
      Tailor-made executive search from start to finish.
    • Payroll + EOR
      Payroll dispersal and EOR across 15+ LATAM countries.
  • Pricing
  • Jobs

0

112
Views
Sampling from multiple tables

We are currently in the process of writing quite complex spark jobs that contain multiple joins and a filters across multiple tables.

We'd like to unit test these jobs with actual data but the real data lives in the cloud (S3 buckets) spreads over dozens of tables (orc files), millions of rows each.

The inherent problem is that sampling from multiple tables can cause joins on the samples to produce no results since there is a possibility that certain IDs don't exist in both sampled tables. Is there a way (a heuristic or a tool) to sample data (for example 1000 rows) from each tables in such a way that joins on foreign keys still provide a useful amount of rows?

That way we can use the sampled data offline to unit test the jobs.

about 4 years ago · Santiago Trujillo
Answer question
Find remote jobs

Discover the new way to find a job!

Top jobs
Top job categories
Business
Post vacancy Pricing Sales
Legal
Terms and conditions Privacy policy
© 2026 PeakU Inc. All Rights Reserved.
Andres GPT
Show me some job opportunities
There's an error!