LLMs Don't Edit Videos: Why the Agent Harness and Tools Matter
Your LLM isn't editing video; tools like FFmpeg are. Understanding the difference between the model, the harness, and the tools is key to building AI agents.
The LLM is not editing your video. ffmpeg is. Every chop, splice, and overlay is a tool call issued by an agent harness, and the model in the loop is doing thought work, not file work. If you've seen someone online claim that WAN, LTX, or any other video model is "editing" footage directly, they're flattening three very different components into one. Here is the architecture, in the order it actually runs.
The Three Components People Keep Merging
When an agent produces a finished video, three things had to happen, and they are not the same thing. In the current hype cycle, marketing departments tend to lump all of this under the banner of "The Model," but that is a fundamental misunderstanding of the engineering stack. To build something that actually works—especially in a niche like fitness tech where precision matters—you have to separate these concerns.
- An LLM: Decides where text should go, what the cuts should be, and what to install. It provides the reasoning capability.
- A Harness: Keeps the goal alive, holds the memory, runs the loops, and dispatches work. This is the persistent state.
- A Set of Tools: Usually ffmpeg plus whatever else is on the local box, actually moved the bytes. These are the muscles.
That last point is the one that breaks most online arguments. The LLM does not have a "chop this clip" button. It has a tool description for ffmpeg, and it emits a command. The shell runs the command. The pixels change. The model never touched the file. It merely suggested a string of text that, when executed by a computer, resulted in a file change.
A Compact Map of Who Does What
| Component | Role in the Stack | What It Actually Produces |
|---|---|---|
| LLM | Decides and plans | Tool calls, installation commands, next-step reasoning |
| Harness | Holds the goal, memory, and loop | The agent's persistence and orchestration |
| Tools (ffmpeg, codecs, scripts, CLIs) | Executes on the local system | The actual file mutations, render, encode, splice |
| APIs / MCPs | Interface to data | Structured reads and writes against external services |
If you remember nothing else: the LLM talks, the harness remembers, and the tools do.
The Heart Analogy, Not the Brain Analogy
The most common mental model is wrong. People call the LLM the "brain" of the agent. It is not. The LLM is closer to a heart. Without it, blood doesn't flow. With it, everything else can move. But a heart does not have goals. It does not have a plan. It pumps.
This matters because the analogy is not just poetic. The capacity of the model sets a ceiling on how fast the whole agent can move, the same way cardiac output caps how much work the rest of the body can sustain. If the LLM is slow to produce the next token, the harness stalls, and the tools sit idle. A local model running on a single machine, on a small parameter count, will not pump the same throughput as a frontier cloud model on high-performance inference hardware.
On my own setup, the model is distributed across three machines, with the GX10 handling the compute-heavy prefill phase before fanning out the generation. Even a small model runs very fast in this configuration. However, if I tried to load a heavier local model—like GLM 4.7 or Minimax 25—into the same memory pool, it would either fail to load or crawl at a few tokens per second. That is the heart not pumping enough. The tools and harness can be perfect, but if the pump is weak, the agent will feel unresponsive and clumsy.
Why the Harness is the Brain
One of the founding researchers of LLMs made the cleanest version of this argument, and it is the part I think most people skip past. An LLM, by itself, cannot have a goal. Next-token prediction is not a goal; it is a statistical probability. A long-form, persistent, multi-hour objective is not something that emerges from a single transformer forward pass. The LLM has no inherent concept of "I want to finish this edit.".
The harness does. A harness (like OpenClaw or custom Python wrappers) can be given a goal, retain that goal in a database or context window, hold a memory structure, run loops, recover from errors, and keep the agent pointed at the same target across many individual tool calls. That is brain-like behavior, and it lives in the harness code, not in the weights of the model.
The companies that try to make the LLM itself the brain will fail. I don't think Anthropic or OpenAI believe the model is the brain internally. The mistake lives in the user community, where the LLM feels so humanlike in conversation that we project the rest of cognition onto it. But conversation is the easiest thing for it. Long-horizon planning against an external world is a different beast entirely.
What the Tools Actually Look Like in the Wild
When you ask an agent to navigate a website or edit a video, the harness doesn't usually have a built-in browser or video encoder. Instead, the agent might take a screenshot of the current state, read the pixels back through a multimodal model, reason about where the cursor should go, and then issue mouse-move and click calls to the operating system. Every one of those steps is a tool call against the local system.
The same is true for video. If you point an agent at a stack like Fable or Codecs, it is going to pull down dependencies to your machine. It installs Git repos, codecs, encoder binaries, and sometimes a full model checkpoint. The video does not appear because a chatbot typed it into existence. The video appears because the harness decided to call ffmpeg with a complex string of arguments—filters, bitrates, and timestamps—and ffmpeg did the heavy lifting of crunching the pixels.
The Fitness Tech Connection: Speediance and Beyond
The first time this clicked for me was when I pointed a free-tier Google Gemini subscription at OpenClaw very early on. I asked it to build the first version of my fitness tracking system. I did not write a single line of code. The harness decided what to install, and the LLM, through the harness, started issuing tool calls: it cloned GitHub repositories, pulled down Garmin Connect's API client, wired up the data flow, and installed the missing libraries on my local box.
None of that was "an LLM editing a file." That was a harness driving a planning loop, a model generating the right next command, and a local toolchain executing it. This has two major takeaways for fitness technology:
- The local system is part of the product: If you block installs or starve the machine of memory, the agent fails. The agent is a full-stack entity, not just a window into a cloud chatbot.
- The API is the frontier: Tools do file work, but APIs and MCPs (Model Context Protocol) do data work. Speediance is launching an API later this year for their equipment data, which is a massive shift. Once there is a clean API for gym hardware, the harness can loop in a workout, pull the set data, and reason about it. The LLM still doesn't "edit" your workout; the agent calls the API to update your records based on reasoning.
Choosing the Right Mental Model
When you see a claim like "this LLM edits video," substitute the actual components and see if the claim still holds. If the LLM is doing the editing, which tool is it calling? If it is calling ffmpeg, then ffmpeg is editing the video. If there is no tool, there is no edit—there is only a text-based description of what an edit could look like.
The cleanest version of the argument is also the shortest: video models are not directly editing files; FFmpeg and other tool calls are doing the work, and the harness gives the agent goals, memory, loops, and practical capability. Once you hold those three pieces apart, the rest of the AI-agent stack, including the parts that will eventually touch your training data, becomes much easier to reason about.
This is a topic-edited cut from the longer livestream. The full argument, including the screen-capture walkthrough of the GX10 setup, is on YouTube: LLMs Do Not Edit Videos. Tools and Harnesses Do.
Go deeper
Follow the AgentStack system from build notes to daily episodes
The AI work has its own path now: implementation reports, daily analysis, and the podcast archive.
Keep Reading
Related Posts

Switching From OpenClaw to Hermes Agents: My Current AI Workflow

Building BJJ Buddy Live: A Speediance Workout and AI Dev Session
