Corpus
The versioned corpus separates manual captions, automatic captions, videos without captions, multiple languages, Shorts, long videos and known unavailable cases. Fixtures must be public, stable enough to rerun and documented without storing complete transcript text.
Execution
Cold and warm requests are measured separately. A cold run clears only the test cache entry. A warm run requests the same canonical video and language again. Generated jobs are followed until completion or a terminal error. The runner records the Capslane image tag, worker version, extractor versions and selected model.
Metrics
The report includes native success rate, generated success rate, end-to-end success rate, p50 and p95 latency, cache rate, error distribution and cost observed by path. Timestamp quality is evaluated on a manually reviewed subset rather than inferred from segment count.
Reporting rules
Every number must refer to a dated run and a corpus version. Removed videos remain visible as unavailable fixtures instead of being silently replaced. Competitor comparisons use the same URLs, options, region and time window. Self-reported competitor statistics are not mixed with measured results.
Limits
YouTube availability varies by region, age, account state and upstream changes. A benchmark describes the observed environment and does not guarantee future extraction. The public dataset contains identifiers and aggregate results, not copyrighted transcripts.