YouTube clone script
Not a one-click installer or a packaged product — a small, readable codebase you fork and own. Every route is traceable from the browser form down to the R2 key it writes, which is the point: you are meant to change it.
The routes
| Route | What it does |
|---|---|
GET / | Server-rendered listing plus the upload form |
GET /watch/<id> | Watch page; unknown ids return a real 404 |
POST /api/videos | Creates metadata, returns an id and a delete token |
PUT /api/videos/<id>/media | Streams the video bytes into R2 |
PUT /api/videos/<id>/thumb | Uploads the thumbnail |
GET /api/media/videos/<id> | Read proxy with Range support → 206 |
DELETE /api/videos/<id>?token=… | Deletes the D1 row and both R2 objects |
/robots.txt, /sitemap.xml | Generated per request, URLs follow the host |
Why upload is two requests
The obvious design is one multipart/form-data POST carrying the file. It is also the one that fails: parsing multipart pulls the whole body into the Worker, and the Worker has a 128 MB memory ceiling — before you even get to the request body cap.
So upload is split. The first call sends only JSON — title, description, content type, size, duration — and returns an id plus a delete token. The second PUTs the raw bytes to that id. Because the second endpoint never parses the body, it can pipe it straight through to R2:
await env.MEDIA.put(key, request.body, { httpMetadata });Workers charge for CPU time, not wall time, and waiting on IO is not CPU. A request that spends two seconds streaming 50 MB into a bucket costs roughly the same CPU as one that returns instantly. Verified on the free plan: a 50 MB PUT lands complete, well inside the 10 ms per-invocation budget.
The split has a second benefit: metadata is committed before the bytes arrive, so a failed upload leaves a row you can delete rather than a row you never knew about.
Bindings
Two bindings, both declared in wrangler.toml:
[[r2_buckets]]
binding = "MEDIA"
bucket_name = "tubeforge-media"
[[d1_databases]]
binding = "DB"
database_name = "tubeforge-db"
database_id = "<your-id>"One bucket holds both kinds of object, separated by prefix — videos/<id>and thumbs/<id>.jpg. Keeping them in one bucket means one binding and one set of lifecycle rules, and the read proxy enforces the prefixes so it cannot be used to browse arbitrary keys.
Deploy it
- Create the bucket:
npx wrangler r2 bucket create tubeforge-media - Create the database —
npx wrangler d1 create tubeforge-db— and paste the returneddatabase_idintowrangler.toml. - Create the table:
npx wrangler d1 execute tubeforge-db --file migrations/0001_create_videos.sql - Change the domains in the
deployscript to your own; it is currently hardcoded to--domains tubeforge.dev www.tubeforge.dev. npm install && npm run deploy
Wrangler needs to be authenticated first (npx wrangler login). The deploy script binds two hosts on purpose: apex and www, with www redirecting to apex so search engines see one canonical site.
What you will want to change first
- The 95 MB upload ceiling and the 9 GB total quota, both in
src/lib/limits.ts. The quota exists because there is no user system — total size is the only thing stopping the free 10 GB being filled by strangers. - Authentication. Deletion is guarded by a token stored in the uploader's browser and nothing else. Fine for a demo, not for anything with accounts.
- Transcoding — the real gap. TubeForge stores exactly what you upload and never re-encodes. A production video site needs a transcode pipeline producing multiple bitrates, which is the single largest cost and complexity item this project deliberately leaves out. If you are comparing against a full YouTube clone, this is the difference that matters.
Where to go next
What a YouTube clone is made ofWhat it costs to runSee it working