AI · Code · Design · Film · 2026 · In dev
EditOne
A desktop video editor where you and an AI assistant share one timeline. Every edit stays visible, undoable and yours to finish.
- Role
- Product design + engineering
- Status
- In development · A–W built
- Built
- October 3–6, 2026
- Stack
- Electron · React · TypeScript · WebGL2 · Claude

Note
Drop footage. Say what you want. Get an edit you can work with.
EditOne is a desktop video editor where a person and an AI assistant (Claude) share one timeline and one set of tools. You can cut by hand, ask for changes in chat, or hand over a one-line brief and get a finished, beat-synced first cut. Whatever the assistant does lands as ordinary, undoable timeline edits you can keep working on. It is a full editor: a GPU compositor, color grading, titles, captions, audio mixing and export. When a tool doesn't exist yet, the assistant can write a new GPU effect or transition, test it and add it to the library.
Brief
The problem and the idea
Most AI video tools hand you a finished render: if one cut is wrong, you regenerate the whole thing. Traditional editors give you full control but leave all the work to you. EditOne joins the two. The AI is a collaborator that edits the same project you do, not a black box that replaces the edit.
Three principles shaped every decision:
- The AI edits the way you do. Every assistant action goes through the same command layer as a manual edit, so it is visible, undoable and editable. The AI never produces a baked video.
- The director decides; the engine executes. The model makes creative choices and writes an edit plan. A deterministic assembler does the frame-exact work: beat snapping, trims, handles and loudness.
- Every feature has three faces: a UI control, an operation the assistant can call, and a test. A feature isn't done until the AI can use it too.
The work
What it does
One project, three ways to work. You can switch between them at any moment.
| Mode | You do | EditOne does |
|---|---|---|
| Manual | Edit on a magnetic multitrack timeline: blade, trim, roll, slip, slide, keyframes, masks | Real-time GPU preview, an Inspector for every property, full undo history |
| Collaborative | Ask in plain words: "match the second shot to the first", "cut the ums", "make the title type on" | Plans and applies the change with the same tools, checks its own frames, explains what it did |
| Fully automatic | Pick a recipe (montage, UGC ad, explainer…), a length, an aspect ratio and a vibe | Analyzes the footage, writes an edit plan, builds a beat-synced cut, reviews it and reports what it checked |
A full editor underneath. About 40 GPU transitions, an effects stack with masks and adjustment layers, color correction and grading with looks, LUTs, shot matching and scopes. There is a text engine with animated titles, captions edited by transcript, and an audio toolkit with cleanup, ducking, loudness and sound effects. Layout templates and smart reframe turn 16:9 into vertical, and film styles range from Archival 16mm to VHS.
Extensible on the fly. Ask for "a heat-haze shimmer" or "the second shot bleeds in from the corners like ink", and the assistant writes a new GLSL shader. EditOne lints it, compiles it, test-renders it and saves it to the project library with live previews and editable parameters. People can write their own in the built-in shader editor.
Connected providers. With your own keys, the assistant can narrate with ElevenLabs voices (with word timings), pull licensed stock from Pexels and Pixabay, and generate new shots with fal. Every paid call shows a cost card you approve first. Keys live in the main process and never reach the UI.
The build
How it works
Every edit, human or AI, is a pure command on one project model, so undo, autosave and the assistant all see the same history.
- One command layer. Timeline positions are integer frames; edits are pure functions producing Immer patches. The assistant's operations compile to the same commands and run inside one atomic transaction per turn, so "Undo this turn" takes back a whole AI build in one step.
- Director, then Assembler. For a full build, Claude writes an Edit Plan (JSON): sections, shot picks, music, look, text. A deterministic Assembler turns it into frame-exact clips with a beam search over shot costs, snaps cuts to the beat and lands the ending on the music's last hit. Creative judgment stays with the model; precision stays in code.
- Media understanding. On import, footage is analyzed locally: shot detection, whisper.cpp transcripts (Vulkan GPU build), a custom beat tracker, face detection and voice activity (ONNX), and text embeddings for search. Claude adds vision summaries of each shot.
- Preview equals export. One WebGL2 compositor (GLSL, float render targets) draws both the viewer and the exporter, fed by WebCodecs hardware decode. Exports encode with WebCodecs and mux with Mediabunny; ffmpeg handles probing, proxies and loudness.
- An agent that checks itself. Tools stream from Claude and every input is validated with Zod before it runs. A bad call applies nothing and comes back with the failing operation and a hint, so the model corrects itself mid-turn. After grading, titling or building, it looks at rendered frames and lint results before reporting.
- Security boundaries. The Electron main process owns the network, API keys (OS-encrypted) and providers. The sandboxed renderer only talks to it through a typed IPC contract.
Manual edits and the assistant's edits meet in the same command layer; the viewer and the exporter draw from the same compositor.
Results
Craft and rigor
Each build step had written done-criteria, and each was signed off with a measured result, not a demo.
- Golden frames. Over 300 reference frames (every transition, effect, look, text animation and film style) re-render and are compared pixel-for-pixel, so a visual regression fails the build.
- Preview equals export, measured. Exported frames are compared against the viewer's frames by SSIM, including through custom shaders and effect stacks.
- Replayable AI. Live Claude sessions are recorded once and replayed in tests, with deterministic IDs. Agent end-to-end tests run on real model output for free, and changing a prompt means re-recording.
- Property tests. 1,000 random command sequences, then undo all, must return the exact starting project.
- Coverage. 63 unit test files and 38 Electron end-to-end specs, across about 53,000 lines of TypeScript.
| Area | Measured result |
|---|---|
| Auto build | 30 s vertical montage from a brief in 142 s; 23 shots, 100% of cuts on the beat; $0.60 of API use |
| Assistant | 10 live editing prompts, all applied as asked; 32 requests, $0.26 |
| Export | 3-minute 1080p30 timeline in 21.2 s (8.5× realtime), −14.0 LUFS, A/V within 2 ms |
| Playback | 10 minutes of 4-layer 1080p30: 0 dropped frames, A/V within 1 frame |
| Timeline | 500 clips dragged at once: 16.8 ms p95, 0 long frames |
| Transcripts | Word error rate 1.7%; local whisper at 37× realtime on an RTX 3080 |
| Effects | 5-effect stack at 1080p: 1.5 ms p95 per frame |
| Reframe | Face kept in frame 100% converting 16:9 to 9:16 |
Devlog
Role, stack and timeline
I designed and built EditOne: the product plan, the UX, the architecture, the rendering and media engines, and the AI system.
- Timeline: planned as 26 build steps (A–Z) with written done-criteria. Steps A–W were built and signed off between Oct 3 and Oct 6, 2026, over 114 commits. Still to come: reference matching and variations, a full recipe library, then polish and release.
- App: Electron, React 19, TypeScript (strict), Vite, Tailwind v4, Radix, Zustand + Immer.
- Media: a custom WebGL2/GLSL compositor, WebCodecs, Mediabunny, an LGPL ffmpeg build, whisper.cpp (Vulkan), ONNX Runtime.
- AI: Claude (Opus 5.5) through the Anthropic SDK: streaming tool use with Zod validation, prompt caching, adaptive thinking.
- Providers: ElevenLabs, fal, Pexels, Pixabay.
- Testing: Vitest, fast-check, Playwright for Electron, golden-frame and parity harnesses.
Strategy
What I learned and what's next
- Give the AI the user's tools, not its own. Routing every assistant action through the same commands as a click made AI edits trustworthy: visible, undoable and editable. It also meant every feature had to be designed for both a person and a model.
- Split judgment from precision. Letting the model plan and code execute gave cuts that are 100% on the beat and builds that are repeatable, which a model working frame by frame could not guarantee.
- Let the model look at its work. Showing the assistant rendered frames after each change caught things no schema could: a subtitle lost against a striped shirt, a washed-out shot match, a black bar on a generated clip.
- Make AI testable. Recording live sessions and replaying them turned a nondeterministic system into one that runs in CI.
Next: matching the pacing and color of a reference edit, versions and variations, a full recipe library with a stronger self-review, then performance work and a packaged release.