Skip to content
Carla Prados

Blog / Research

Three sources, one frame rate, and a 35× disagreement

Carla Prados Bodega 14 Aug 2026 · 2 min read

Photo by Markus Spiske on Unsplash

Every video in my dataset reported its frame rate three times, and the three readings did not agree. Resolving that was most of the project.

The number that was not a number

I had 181 videos and a pipeline to write. The plan was simple: read each file, track the marker, convert pixel displacement to millimetres and time to seconds. Frame rate is the conversion factor for the second half of that sentence, and I assumed it was a fact I could look up.

It is not one fact. It is three.

OpenCV reports a frame rate. ffprobe reports r_frame_rate. ffprobe also reports avg_frame_rate, which is a different thing, and the filename convention we used encodes an intended value that is a fourth claim of its own.

How wrong it can go

On one file, two of those readings imply a clip of 983 seconds. The others imply 28.3 seconds. Same file.

That is roughly a 35× difference in derived timing — and timing is the denominator of every velocity in the dataset. A pipeline that silently picks whichever value it read first will produce numbers that look completely reasonable and are wrong by more than an order of magnitude.

A value you did not choose deliberately is not data. It is a default that survived.

The disagreement was not an edge case in a handful of files. Once I actually checked, the conflict was present in all 181.

What I did instead

Three things, none of them clever:

  • Record every claim. The inventory stores all sources for each file rather than collapsing them on read.
  • Resolve one authoritative value, by an explicit rule, and write down which rule fired.
  • Keep the conflict visible. Each record carries an inconsistencies field, and the analysis summary states the distinction in words, so a reader downstream cannot accidentally reinterpret the timing.

That last one matters most. Resolving the conflict is easy; the risk is that the resolution becomes invisible, and six months later somebody reasonably assumes the number was never in question.

The part I would tell my past self

I lost about a week to this and initially resented it, because it produced no findings. That was the wrong way to count it.

The campaign’s actual results — a curvature-to-asymmetry relationship, a systematic speed overshoot — are only worth stating because the timing underneath them is defensible. The week was not a detour from the analysis. It was the analysis acquiring the right to exist.

#data-validation #metadata #python