Metrics¶
This page describes the metrics the toolkit computes, grouped by analysis layer. The names are the field names in the records and the column names in the CSV files.
Bitstream statistics¶
The bitstream statistics come from videoparser-ng. They are written as one frame_bitstream record per frame, in presentation order, for MPEG-½ video, H.264, HEVC, VP9 and AV1.
Please refer to the public repository for more information.
Motion vector statistics
The toolkit uses the "legacy" build of videoparser-ng, which P.1204.3 requires. For H.264, HEVC and VP9, the motion vector statistics are normalized by the temporal distance to the reference frame. They therefore differ from those of a standard videoparser-ng build. All other values are identical.
Pixel metrics¶
The pixel metrics come from AVEQ's video analyzer, the same analyzer as in the Video Analyzer SDK. They are no-reference metrics: they are computed from the decoded pictures alone, without a reference. They are written as one frame_pixels record per analyzed frame.
The metrics are grouped into features, which you select with --pixel-features. Features marked with "default" are computed when you do not set the option. Metrics of features that were not computed are left out of the records.
| Feature | Default | Metrics | Range and meaning |
|---|---|---|---|
blockiness |
yes | blockiness |
0 (none) to 1 (severe). Multi-scale detection of block edges. As a rough guide: 0.2 minimal, 0.4 noticeable, 0.6 significant. |
blockiness_ntia, blockiness_vp9 |
Blockiness by the NTIA and VP9 methods, unscaled | ||
blurriness |
yes | blurriness |
0 (sharp) to 1 (blurry), scaled to 1080p |
blurriness_raw |
Blur magnitude before the mapping to 0–1 | ||
brightness |
yes | brightness |
Average gray value, 0 to 1 |
blackscreen |
yes | black_screen |
Whether the frame is completely black |
spatialcomplexity |
yes | spatial_information |
Spatial information (SI) after ITU-T P.910 (legacy definition) |
noise |
yes | noise |
Noise estimate on the luma plane |
temporalchange |
yes | temporal_information |
Temporal information (TI) after ITU-T P.910 (legacy definition) |
freezing |
Whether the frame repeats the previous one | ||
scene_change, scene_score |
Whether a scene change starts at this frame, and its score | ||
sad_normalized |
Difference to the previous frame, 0 to 1 | ||
jerkiness |
yes | jerkiness |
Jerkiness from repeated frames in a group, 0 to 1 |
chroma |
no | chroma |
Colorfulness, 0 (gray) to 1 |
contrast |
no | contrast, contrast_michelson |
Contrast from the 1st and 99th percentile, and Michelson contrast, 0 to 1 |
radialprofile |
no | sum_of_high_frequencies |
Sum of high frequencies of the radial frequency profile |
autoenhancement |
no | white_level, black_level |
Irregular white and black levels, 0 to 1, higher is worse |
finedetail |
no | fine_detail |
Fine detail, 0 to 1, higher is worse |
--pixel-features all computes all features, which takes about twice the CPU time of the default set. Each frame_pixels record also contains the width and height of the analyzed picture.
P.1204.3 quality score¶
ITU-T P.1204.3 is a standardized bitstream-based model that estimates the quality a viewer perceives, as a MOS from 1 (bad) to 5 (excellent). It uses the bitstream statistics (QP, motion, frame sizes and types). See ITU-T P.1204 for background on the model.
The toolkit writes model_score records with model set to p1204.3 in three scopes:
per_second: one score per second of media timeper_gop: one score per group of picturesoverall: one score for the whole stream
Each record contains the start and end of the scored time range, the device (pc or mobile) and the display resolution the model was set up for (--device and --display), and in details the intermediate values of the model.
When the model cannot produce a score, the record has no score, but a reason with a code:
unsupported_codec: the codec is not covered by P.1204.3 (for example MPEG-2 or AV1)insufficient_data: the stream is too shortmodel_error: the model failed
Full-reference metrics¶
compare computes the full-reference metrics with FFmpeg's libvmaf, psnr and ssim filters. The values are identical to what you get from FFmpeg with the same settings. They are written as one frame_fullref record per frame of each encode.
VMAF (vmaf):
score: VMAF score of the frame, usually 0 to 100model,model_version: the model used, for examplevmaf_v1.0.16_3d0hfeatures: the elementary features of the model under stable names (for exampleadm3,motion3,cambi)
PSNR (psnr):
y,u,v: PSNR of the luma and the two chroma planes in dBavg: PSNR over all planes in dBmse_y,mse_u,mse_v,mse_avg: the mean squared errors
When a plane is identical to the reference, its MSE is 0 and the PSNR is infinite. The PSNR value is then left out of the record.
SSIM (ssim):
y,u,v: SSIM of each planeall: SSIM over all planes, andall_dbthe same in dB
Each frame_fullref record also has a reference field with the reference frame it was compared with, and a mapping field that says how the frame was matched:
exact: the timestamps match after applying the offsetnearest: the nearest reference frame in time was used (different frame rates or timestamp jitter)repeated: the same reference frame as for the previous frame
Summary statistics¶
For every metric, the toolkit writes summary records with descriptive statistics over the whole stream:
count: number of valuesmean,std: mean and population standard deviationmin,maxp1,p5,p50,p95: 1st, 5th, 50th (median) and 95th percentileharmonic_mean: for VMAF only
record_type and metric name the field the statistics are computed from, for example frame_bitstream and qp_avg. For bitstream metrics, there are additional summaries per frame type, with frame_type set to I, P or B. For full-reference metrics, reference gives the ID of the reference input.
In compare, comparison records give the difference of each summary between an encode and the reference (delta = encode minus reference).