Proposed additive publication plan
This is a review proposal, not authorization to deploy. The destination remains
the existing openai-film-fieldnotes-100 Cloudflare Pages project. No Gsites,
replacement project, public file host or R2 bucket is part of this plan.
Keep the benchmark intact
Proposed new route: /videos/editing/benchmark/library/. Keep the approved
five-treatment benchmark at its current route, with its players, evidence and
downloads unchanged. Add only an approved navigation link to the new library.
Do not turn the original benchmark into a redirect or replace its files with
the new engine.
The saved task-start live manifest has 4,962 entries. The local public-build manifest has only 3,886, so it cannot be used as the deployment baseline. Before any approved deployment, retrieve the latest live release and manifest again, reconcile concurrent additions, and create an additive staging directory. Fail staging on an unexplained old-file deletion or byte change. Preserve the pre-deploy live manifest and the exact approved navigation delta.
Run a publication-only privacy scan over code, catalogs, documentation, manifests and evidence. Local provenance currently includes machine-specific executable paths and diagnostic logs; do not copy those unchecked. Keep original local records intact. If a public derivative redacts nonbinding diagnostic metadata, record both hashes and the exact removed fields, and prove that code/input/asset bindings, inspected pixels and source mapping are unchanged. Never edit a binding, regenerate a hash and describe it as the originally reviewed record.
Budget the evidence rather than dropping it
The conservative target is the Free-plan ceiling: 20,000 files and 25 MiB per
asset, checked September 5, 2026 against
https://developers.cloudflare.com/pages/platform/limits/.
Do not assume a paid upgrade or use another host to evade those constraints.
Raw native PNGs from 200 studies must not be copied wholesale into the site. Retain all native frames in downloadable, checksum-indexed evidence archives, split below 24 MiB per part. Keep source comparisons, original/adapted MP4s, small contact sheets, selected recipe JSON and a compact index directly addressable. Publish the runnable core separately from per-recipe asset packs so an agent retrieves only the content it needs.
The published gallery must distinguish a directly viewable contact sheet from a full-resolution frame in an evidence download. Do not leave broken native-PNG URLs after packaging. Each packed frame keeps its source/local index, SHA-256, archive part and member path. Package manifests must bind the actual shipped code, inputs and assets, not a later rebuild.
These packaging and download paths are proposed, not yet a verified 200-case deployment. A final file-budget check must examine the actual complete staging directory, including the preserved release, not an average-size projection.
First-batch size measurement
evidence/publication-capacity-first20-review02.json pins the selected local
artifacts measured on September 5, 2026. The 20 draft selections contain 7,217
measured files / 1,992,382,589 bytes, including 6,486 native PNGs and 731 directly
served or editable-asset candidates. No measured direct file exceeds 25 MiB.
All selected manifest asset hashes match, but this sizing run does not reassess
render currency, permissions or fidelity.
A conservative stored-ZIP calculation estimates 78 native-frame parts below 24 MiB, reducing that measured subset to 809 served files. Linear extrapolation to 200 treatments, the task-start 4,962-file baseline and a chosen 500-file shared reserve totals 13,552 files. This is a capacity estimate, not 200 actual examples or a staging pass. No archives were built, and new treatments may be larger. Final packaging must include the actual docs, licenses, redacted reports, download indexes and extraction helpers; the reserve does not prove they fit.
Reproduce into a new, non-overwritten evidence path with
python3 scripts/estimate_publication_capacity.py --out evidence/NEW-NAME.json.
Five archive-budget unit tests pass in evidence/capacity-estimate-tests.log;
archive extraction, public URLs and the complete real staging budget remain
unverified.
Verified local packaging prototype
The later review03 run creates real archives for all 20 selected drafts:
evidence/native-packages-first20-review03/package-report.json records 6,582
native PNGs in 78 deterministic stored-ZIP parts. Their total size is
1,755,620,179 bytes; the largest is 25,151,162 bytes, strictly below 24 MiB.
Every member is read back, hash-checked and decoded. Each per-recipe index
preserves source attribution, rational timing and the exact frame/member map.
This actual run does not replace or retroactively validate the older capacity
estimate, which used a different selected snapshot.
Use scripts/package_evidence.py --recipe ID --out evidence/NEW-DIRECTORY
through the documented Python research environment. It refuses existing
destinations, unsafe paths, missing or changed inputs, oversized members and
stale render/comparison bindings. Source frames remain research-only; retain
their attributed index with the archives. The archive unit suite has 30 cases;
the actual run also invokes the delivered render verifier.
This prototype does not package the full runnable toolkit or editable-asset downloads, rewrite public frame URLs, create privacy-safe public report derivatives, or prove a complete 200-treatment staging budget. Those steps and the live checks below remain required after approval. Nothing was deployed.
Approval and live checks
The first review approves or revises coverage, the reuse contract and the first 20 source choices/examples. It does not waive unresolved fidelity findings and does not authorize publication. Complete the source reviews and unseen reuse trials before claiming the requested library outcome.
Obtain explicit publication approval for the frozen addition. Deploy only that staging directory to the named existing project. Then verify every catalog and download URL, byte hash and range-capable video response; play original, adaptation and synchronized comparisons; exercise search, filters, deep links, pagination and mobile layouts; and compare all preserved release files against the pre-deploy manifest. Keep failed checks visible and do not report a partial live deployment as completion.