Business
Jobs
  • About Us
  • Solutions
    • Job Postings
      Post your job and receive qualified candidates in 48h.
    • Candidate Assessments
      500+ technical and psychological tests, plus anti-fraud.
    • Headhunting
      Tailor-made executive search from start to finish.
    • Payroll + EOR
      Payroll dispersal and EOR across 15+ LATAM countries.
  • Pricing
  • Jobs

0

247
Views
How can HTML5-Video record the PTS code or Number of a paused frame to allow a subsequent viewer to jump to that exact frame?

I have read this post that answers with 4 commercial players that implement frame-accurate seeking and I will investigate them, but we are trying to build a tool for labeling objects in a video and we need to know precisely the frame that they have labeled so a subsequent watcher can validate their label.

We ask the users to watch until they see our desired object. Then pause the video and navigate to the "Best Frame". We give them buttons to go back and forward (1 second/Average FPS) to go move one frame at a time (unless our video has dropped frames, then they must push those buttons once for each dropped frame.)

[Bonus question: can we tell if gave us a new frame or just the same frame with an updated time-in-video? If we got the same frame, we could loop until the first new frame so the user wouldn't have to keep pressing the key!]

When they arrive at the best frame, we can record the time of that frame, but we are finding that the time often comes back with different milliseconds each time someone tags the exact same frame. So we can't reliably have a subsequent user seek the exact same frame to validate the first watcher's label. We are using Canvas to allow the first user to draw a bounding box around our object of interest. Yet we can't reliably redraw the bounding box if it's on the wrong frame as our objects move quickly.

Ideally, if it can't be a sequential frame number, can it be PTS code? Is there any other way we can determine the precise frame each time? We are building this for a controlled audience that we can train, so creative solutions are acceptable.

We are open to a "cheat" because we control each video in our workflow. If we can write "something" into the file that can be read by the HTML-Video, we could seek to the approximate time minus 0.5 seconds and then read the subsequent frames for "something" that can tell us we have an exact match.

about 4 years ago · Juan Pablo Isaza
Answer question
Find remote jobs

Discover the new way to find a job!

Top jobs
Top job categories
Business
Post vacancy Pricing Sales
Legal
Terms and conditions Privacy policy
© 2026 PeakU Inc. All Rights Reserved.
Andres GPT
Show me some job opportunities
There's an error!