Skip to content

Aligning Inputs

Full-reference metrics are only meaningful when each frame of an encode is compared with the matching frame of the reference. If the encodes do not align due to some processing, we do our best to find the best match for each frame.

Alignment methods

--align selects how compare matches frames:

  • pts (default): match frames by their presentation time, counted from the first frame of each input, plus the offset given with --offset. Use it when encode and reference start at the same picture, or when you know the offset.
  • auto: find the offset automatically, then match by presentation time. Use it when you do not know the offset.
  • none: match frames by their index (first frame with first frame, and so on), ignoring timestamps. Use it only for inputs with broken timestamps.

When the frame rates differ, each frame of the encode is mapped to the reference frame nearest in time. Reference frames without a matching encode frame are counted as dropped, and encode frames mapped to the same reference frame as the previous one as repeated.

Visual alignment

Future versions may add a visual alignment tool to check the frame matching. For now, we recommend using tools like video-offset-finder for this purpose.

Finding the offset automatically

For example, a recording of a channel starts about 2 seconds later in the content than the reference:

surfmeter-media compare --align auto -o results reference.mkv recording.ts

The toolkit compares small thumbnails of the luma plane of both inputs at every candidate offset and picks the one that matches best. By default, it searches 5 seconds in each direction. Widen the search with --search-range, or give an estimate with --offset to search around it:

# Search ±15 seconds
surfmeter-media compare --align auto --search-range 15 -o results reference.mkv recording.ts

# Search ±2 seconds around an estimated offset of 30 seconds
surfmeter-media compare --align auto --offset 30 --search-range 2 -o results reference.mkv recording.ts

The offset found is printed in the Offset column of the result table, and stored in the alignment record:

jq -c 'select(.type == "alignment") | {input, method, relative_offset, offset_frames, confidence, matched_frames, dropped_frames, repeated_frames}' \
  results/records.ndjson

confidence ranges from 0 to 1. It is low when the content is static or repeats itself (for example a still image, a slate, or a loop), because many offsets then match about equally well. In this case, check the offset manually, for example in the terminal view, and set it with --offset.

Setting the offset manually

If you know the offset, give it in seconds with --offset. A positive offset means that the encode starts later in the content than the reference:

# The encode starts 1.2 seconds into the reference
surfmeter-media compare --offset 1.2 -o results reference.mkv encode.mp4

# The encode starts 0.48 seconds (12 frames at 25 fps) before the reference
surfmeter-media compare --offset -0.48 -o results reference.mkv encode.mp4

The same offset applies to all encodes of the run. If several encodes have different offsets, use --align auto, which finds the offset of each encode on its own, or run compare once per encode.

Checking the alignment

A wrong alignment shows clearly in the results:

  • VMAF is low over the whole file, even for an encode at a high bitrate.
  • The number of matched frames is much lower than the number of frames of the encode.

In the terminal view, the VMAF track of a misaligned encode stays low throughout, while it follows the content (dips at difficult scenes) for a correctly aligned one.

Transport stream timestamps

All records use the timestamps of the file itself. Transport streams usually do not start at 0 (for example, the first frame of a recording may have a timestamp of 1.48 s or 9,321.6 s). The pts field of every record keeps this original timestamp, so you can find a frame in other tools such as ffprobe. The alignment record gives the offset on both scales:

  • offset: on the original timestamps (reference_pts = input_pts + offset)
  • relative_offset: relative to the first frame of each input, as you would give it with --offset