Skip to content

Metrics

This page describes the metrics the toolkit computes, grouped by analysis layer. The names are the field names in the records and the column names in the CSV files.

Bitstream statistics

The bitstream statistics come from videoparser-ng. They are written as one frame_bitstream record per frame, in presentation order, for MPEG-½ video, H.264, HEVC, VP9 and AV1.

Please refer to the public repository for more information.

Motion vector statistics

The toolkit uses the "legacy" build of videoparser-ng, which P.1204.3 requires. For H.264, HEVC and VP9, the motion vector statistics are normalized by the temporal distance to the reference frame. They therefore differ from those of a standard videoparser-ng build. All other values are identical.

Pixel metrics

The pixel metrics come from AVEQ's video analyzer, the same analyzer as in the Video Analyzer SDK. They are no-reference metrics: they are computed from the decoded pictures alone, without a reference. They are written as one frame_pixels record per analyzed frame.

The metrics are grouped into features, which you select with --pixel-features. Features marked with "default" are computed when you do not set the option. Metrics of features that were not computed are left out of the records.

Feature Default Metrics Range and meaning
blockiness yes blockiness 0 (none) to 1 (severe). Multi-scale detection of block edges. As a rough guide: 0.2 minimal, 0.4 noticeable, 0.6 significant.
blockiness_ntia, blockiness_vp9 Blockiness by the NTIA and VP9 methods, unscaled
blurriness yes blurriness 0 (sharp) to 1 (blurry), scaled to 1080p
blurriness_raw Blur magnitude before the mapping to 0–1
brightness yes brightness Average gray value, 0 to 1
blackscreen yes black_screen Whether the frame is completely black
spatialcomplexity yes spatial_information Spatial information (SI) after ITU-T P.910 (legacy definition)
noise yes noise Noise estimate on the luma plane
temporalchange yes temporal_information Temporal information (TI) after ITU-T P.910 (legacy definition)
freezing Whether the frame repeats the previous one
scene_change, scene_score Whether a scene change starts at this frame, and its score
sad_normalized Difference to the previous frame, 0 to 1
jerkiness yes jerkiness Jerkiness from repeated frames in a group, 0 to 1
chroma no chroma Colorfulness, 0 (gray) to 1
contrast no contrast, contrast_michelson Contrast from the 1st and 99th percentile, and Michelson contrast, 0 to 1
radialprofile no sum_of_high_frequencies Sum of high frequencies of the radial frequency profile
autoenhancement no white_level, black_level Irregular white and black levels, 0 to 1, higher is worse
finedetail no fine_detail Fine detail, 0 to 1, higher is worse

--pixel-features all computes all features, which takes about twice the CPU time of the default set. Each frame_pixels record also contains the width and height of the analyzed picture.

P.1204.3 quality score

ITU-T P.1204.3 is a standardized bitstream-based model that estimates the quality a viewer perceives, as a MOS from 1 (bad) to 5 (excellent). It uses the bitstream statistics (QP, motion, frame sizes and types). See ITU-T P.1204 for background on the model.

The toolkit writes model_score records with model set to p1204.3 in three scopes:

  • per_second: one score per second of media time
  • per_gop: one score per group of pictures
  • overall: one score for the whole stream

Each record contains the start and end of the scored time range, the device (pc or mobile) and the display resolution the model was set up for (--device and --display), and in details the intermediate values of the model.

When the model cannot produce a score, the record has no score, but a reason with a code:

  • unsupported_codec: the codec is not covered by P.1204.3 (for example MPEG-2 or AV1)
  • insufficient_data: the stream is too short
  • model_error: the model failed

Full-reference metrics

compare computes the full-reference metrics with FFmpeg's libvmaf, psnr and ssim filters. The values are identical to what you get from FFmpeg with the same settings. They are written as one frame_fullref record per frame of each encode.

VMAF (vmaf):

  • score: VMAF score of the frame, usually 0 to 100
  • model, model_version: the model used, for example vmaf_v1.0.16_3d0h
  • features: the elementary features of the model under stable names (for example adm3, motion3, cambi)

PSNR (psnr):

  • y, u, v: PSNR of the luma and the two chroma planes in dB
  • avg: PSNR over all planes in dB
  • mse_y, mse_u, mse_v, mse_avg: the mean squared errors

When a plane is identical to the reference, its MSE is 0 and the PSNR is infinite. The PSNR value is then left out of the record.

SSIM (ssim):

  • y, u, v: SSIM of each plane
  • all: SSIM over all planes, and all_db the same in dB

Each frame_fullref record also has a reference field with the reference frame it was compared with, and a mapping field that says how the frame was matched:

  • exact: the timestamps match after applying the offset
  • nearest: the nearest reference frame in time was used (different frame rates or timestamp jitter)
  • repeated: the same reference frame as for the previous frame

Summary statistics

For every metric, the toolkit writes summary records with descriptive statistics over the whole stream:

  • count: number of values
  • mean, std: mean and population standard deviation
  • min, max
  • p1, p5, p50, p95: 1st, 5th, 50th (median) and 95th percentile
  • harmonic_mean: for VMAF only

record_type and metric name the field the statistics are computed from, for example frame_bitstream and qp_avg. For bitstream metrics, there are additional summaries per frame type, with frame_type set to I, P or B. For full-reference metrics, reference gives the ID of the reference input.

In compare, comparison records give the difference of each summary between an encode and the reference (delta = encode minus reference).