Claude can watch videos now, and it never presses play
The short version: a free open-source skill called /watch gives Claude Code a video input. It downloads the video with yt-dlp, cuts it into still frames with ffmpeg, pulls a timestamped transcript from the platform’s own captions, and hands Claude the pictures and the text together. Nothing is ever played. A ten minute video arrives as about eighty photographs, which is one roughly every seven and a half seconds, so anything that happens between two of them is invisible. Knowing that is what tells you how to use it: give it a time range, and pick the right detail level.
None of this is hidden. It is all in the skill’s own documentation, written by its author. What that documentation does not do is put the mechanism and its consequences side by side, which is what this is.
What the /watch skill actually does
The skill is bradautomates/claude-video, MIT licensed, built by Brad Bonanno. At the time of writing it has 15,654 stars on GitHub, checked on 17 August 2026. Stars are not installs, and several write-ups of this repo quote a much older figure or describe those stars as downloads, so treat any number you see attached to it with the same suspicion.
Its own summary of itself is the clearest description anyone has written: you do not have a video input, and this skill gives you one. Four things happen when you paste a link.
One, yt-dlp checks the source for native captions, manual or auto-generated, and grabs them if they exist. This is free and instant, and it is why most YouTube videos never touch a transcription API at all.
Two, it downloads the video itself, unless you asked for transcript-only detail and the captions were already enough.
Three, ffmpeg extracts frames as JPEGs. At the two higher detail levels this is scene-aware: it uses ffmpeg’s scene-change selection and only falls back to uniform sampling when the video is effectively static.
Four, Claude reads each frame path as an image and combines the pictures with the timestamped transcript to answer you.
How to install the Claude video skill
Two ways, depending on where you want it. Both are one-liners, and neither needs an account anywhere.
Install it in Claude Code
Add the marketplace, then install the skill from it. This is the path to use if you live in the CLI.
/plugin marketplace add bradautomates/claude-video
/plugin install watch@claude-video
Or install it anywhere else
The same skill runs on other Agent Skills hosts, Codex and Cursor and Copilot and Gemini CLI among them, through the skills CLI.
npx skills add bradautomates/claude-video -g
For Claude on the web, download watch.skill from the repository’s latest release and upload it under Settings, Capabilities, Skills. That path needs code execution turned on.
Let it install yt-dlp and ffmpeg
Both are handled on first run. On macOS the skill installs them with Homebrew; on Linux and Windows it prints the commands for you to run. Nothing else is required to watch a video that has captions.
Decide about a Whisper key, or decide not to
This is optional, and the skill is deliberate about that: a missing key is something it encourages you to fix, not a blocker. It only matters for videos with no captions, and for local files.
If you want it, Groq is the one to pick. It runs whisper-large-v3 and the skill prefers it over OpenAI because it is cheaper and faster; OpenAI’s whisper-1 is the fallback. Keys live in ~/.config/watch/.env, and the skill sends each one only to its own provider.
--no-whisper.How to pick the detail level
This is the one setting worth understanding, because it decides both what Claude can see and what the request costs you. There are four, set per-call with --detail or globally with WATCH_DETAIL in that same config file.
| Level | Frames | Use it when |
|---|---|---|
transcript | None | You only care what was said. Skips the video download entirely when captions exist, so it is by far the fastest and costs zero image tokens. |
efficient | Up to 50, keyframes | You want a rough visual sense of a long video without paying for it. Fast, because keyframes are already sitting in the file. |
balanced | Up to 100, scene-aware | The default, and the right answer nearly always. Picks frames where the picture actually changes rather than on a timer. |
token-burner | Uncapped, scene-aware | Dense on-screen text, or a long video you cannot narrow down. It warns you past 250 frames, and it means the name. |
The detail level sets a ceiling. Underneath it, the skill also targets a frame budget based on how long the video is, and whichever number is lower wins.
Those budgets in full: about 12 to 30 frames under 30 seconds, around 40 up to a minute, 60 up to three minutes, 80 up to ten minutes, and past ten minutes the detail cap spread thinly with a warning printed. The skill’s own recommendation is to keep videos under ten minutes for best accuracy, because coverage scales inversely with duration.
What Claude cannot see in a video
Here is the arithmetic that changes how you use this. Eighty frames across ten minutes is one photograph every seven and a half seconds. Everything in between is simply not in the pile.
So the failure mode is specific and predictable. Fast motion, a quick gesture, a swipe, a UI animation, a frame of a video game, a three-second toast notification: any of it can fall in a gap and Claude will answer as though it never happened, because as far as its inputs are concerned it did not.
--no-dedup exists for judging subtle frame-to-frame motion.The fix is not a bigger detail level, it is a smaller question. Give it a range and it spends the whole budget there.
/watch URL --start 2:15 --end 2:45 what does the cursor click here?
And when you do need to read small text on screen, raise the resolution rather than the frame count, remembering that 1024 roughly quadruples the image tokens per frame.
/watch URL --detail token-burner --resolution 1024
What it costs you in context
Frames are images, and images are the most expensive thing you can put in a context window. Eighty of them at the default 512px width runs roughly 50,000 to 80,000 image tokens depending on aspect ratio. The transcript beside them is a rounding error, a few thousand tokens at most for ten minutes.
That is a large single spend in a session, and it does not go away afterwards, because everything already in a conversation is re-sent with every subsequent message. If you are watching videos inside a working session, the habits in our memo on Claude Code token usage matter more than usual: watch the video, get your answer, then /clear before you carry on, or those eighty photographs ride along on every message for the rest of the day.
Where your video and your data actually go
Worth stating plainly, because “a skill that watches your screen recordings” sounds worse than it is. From the skill’s own security section:
- The video itself is never uploaded to any API. Only the extracted audio clip leaves, and only when native captions are missing and Whisper has not been disabled.
yt-dlpandffmpegrun locally. yt-dlp only ever requests public data, with no login and no session cookies, and it cannot post anything anywhere.- Keys are not shared between providers. A Groq key only goes to Groq, an OpenAI key only to OpenAI, and neither is logged, cached or written to output.
- Nothing persists outside the working directory and
~/.config/watch/.env.
The audio clip is the one thing to be deliberate about. If a recording contains something you would not paste into a third-party API, run it with --no-whisper and work from frames.
Questions people actually ask
Can Claude actually watch a video?
It can answer questions about one, but it never plays it. The /watch skill downloads the video with yt-dlp, uses ffmpeg to cut it into still JPEG frames, pulls a timestamped transcript from the platform's captions, and hands Claude the frames as images alongside the text. Claude is reading a contact sheet and a subtitle file, not watching playback. That distinction is the whole reason the limits below exist.
How many frames does it actually take?
It scales the budget to the length of the video: about 12 to 30 frames for anything under 30 seconds, around 40 for 30 seconds to a minute, 60 for one to three minutes, and about 80 for three to ten minutes. Past ten minutes it spreads the detail mode's cap thinly across the whole runtime and prints a warning. It never samples faster than 2 frames per second, no matter what you ask for.
Do I need an API key?
Only for videos with no captions. The skill prefers native captions, which yt-dlp pulls for free from the source platform, and most YouTube videos have them. When there are none, it extracts a mono 16 kHz audio clip and sends it to Whisper, which needs either a Groq key (preferred, cheaper and faster) or an OpenAI key. Skip the key and those videos come back as frames with no transcript at all.
Does it upload my video anywhere?
No. The skill's own documentation is explicit that the video itself never goes to any API. yt-dlp and ffmpeg both run locally on your machine. The only thing that leaves is the extracted audio clip, and only when native captions are missing and Whisper has not been disabled. It also never logs into any platform, so there are no session cookies involved, and each API key is only ever sent to its own provider.
Why did it miss the thing I asked about?
Almost always because the thing happened between two frames. At about one frame every seven and a half seconds on a ten minute video, a quick gesture, a fast swipe or a UI animation can fall entirely into the gap. The skill also drops frames that look nearly identical to the previous one by default, which is great for held slides and terrible for judging motion. Narrow the range with --start and --end, or pass --no-dedup.
How many tokens does watching a video cost?
Frames dominate the bill. The skill's documentation puts 80 frames at 512px wide at roughly 50,000 to 80,000 image tokens depending on aspect ratio. The transcript is cheap by comparison, a few thousand tokens at most for a ten minute video. Raising --resolution to 1024 roughly quadruples the image tokens per frame, so only do it when you genuinely need to read on-screen text.
Can I use it to debug a screen recording?
That is one of the things it is best at, with one caveat. Screen recordings are mostly static, so the deduplication pass throws away a lot of near-identical frames and spends the budget on the parts that actually changed. For reading small UI text, raise the resolution to 1024 and narrow the time range rather than scanning the whole recording at high detail.
Sources
Want this built for you?
We write these memos because we build this stuff every day. If you want it working in your business instead of sitting on your reading list, that is literally our job.