Smart cutting
Silence, filler words, dead takes and repeats removed — natural pauses preserved.
Cutframe is not a template editor or a pile of AI buttons. It watches the whole take, decides what matters, builds the timeline, renders it server-side, then critiques and re-cuts its own work.
Footage in, finished film out — with your own asset library, sound design, motion graphics and colour handled inside one place.
Multi-hour files, resumable chunked uploads, phone gallery or a pasted link.
Logos, LUTs, overlays, fonts and past exports — reusable across every project.
Licence-free music, SFX and sound design tucked under the edit automatically.
Title cards, lower-thirds, hooks and 3D elements rendered into the cut.
Exposure, balance and film looks with protected skin tones.
1080p by default, 2K/4K on demand, plus 9:16, 1:1 and 16:9 in one run.
Every stage emits real state — a job either ran, is running, or failed. Nothing is simulated in the UI.
Validation, proxies, waveform + frame index.
01Scenes, shots, faces, speakers, quality scoring.
02Story beats, take selection, pacing curve.
03Deterministic timeline JSON, server-rendered.
04Self-critique pass scores the cut, then revises.
05Each capability carries its real build state. Nothing ships labelled as working before the pipeline actually produces it.
Silence, filler words, dead takes and repeats removed — natural pauses preserved.
Speaker-aware transcripts with emphasis and safe-area placement.
Levelling, denoise, speech clarity and music ducking in the render graph.
Hook, setup, main story, payoff and ending derived from the actual content.
No fixed formula: emotional beats breathe, fast content cuts tighter.
9:16, 1:1, 4:5 and 16:9 with subject tracking across the cut.
Track choice from mood and pacing; cuts snapped to detected beats.
Exposure, white balance and contrast with protected skin tones.
Active-speaker detection picks the camera; punch-ins for single-cam.
Clean cuts by default; dissolves and match cuts only when earned.
Only inserted where a shot genuinely needs visual support.
Restrained impacts and ambience tied to editorial emphasis.
The plan is format-agnostic, so the same understanding pass produces a landscape cut, a vertical short and a podcast version without re-analysing the footage.
After a render you don't reopen a timeline — you say what's wrong. The assistant patches the existing editing plan and re-renders only the affected ranges.
After rendering, a separate evaluator inspects the result for awkward cuts, clipped audio, black frames, caption drift and story coherence. If a dimension fails its threshold, the plan is patched and only the affected ranges are re-rendered.
Illustrative rubric from the evaluation spec. Real scores are produced per render once the evaluator runs against a project.
Create AI Edit is always the primary action. Everything else — projects, storage, exports, the optional manual timeline — sits behind it.
New AI Edit, recent projects, processing, completed, drafts, exports and storage in one view.
Thumbnail, duration, live status, created and last-edited dates, export and quick actions.
Video, audio, caption, B-roll, music, SFX and effect tracks — with “Re-edit with AI” always present.
Object storage for originals, proxies, thumbnails and renders. Databases hold references, never media.
Training pairs raw footage with a professional final edit, its timeline and the decisions behind it — not tutorials.
Cuts, shots, scenes, continuity, rhythm, transitions, audio, composition.
Best-take selection, pacing, reaction shots, B-roll, music timing, colour.
YouTube, Shorts, vlogs, podcasts, gaming, education, documentary, ads, interviews.
What matters, what's boring, what's emotional, what to emphasise or cut.
The edit is scored by a separate evaluator, then revised and re-scored.
Your footage is never used to improve models without explicit, revocable consent.
Signed URLs, scoped access, validated uploads, rate-limited APIs and server-held secrets.
Choose how long media is kept, delete a project and its derivatives at any time.
Anything not yet working is labelled Research or In build. No fake buttons, no fake progress.
Architecture is decided from documented, currently available technology — speech recognition, vision and scene models, FFmpeg-based server rendering, GPU queues — not from invented APIs. Capabilities that need training data or infrastructure we do not yet have are recorded as open work.
Speech, vision, embeddings and a reasoning layer behind swappable providers.
FFmpeg graphs, WebCodecs proxies, GPU queues and cloud rendering.
Film grammar, shot selection, pacing, highlight detection.
Raw + professional-edit pairs with decision-level annotation.
CapCut, Premiere, Descript, VEED, OpusClip, Runway, Kapwing — what they automate vs. leave manual.
Story, pacing, audio, visual, caption and overall professional scores, 0–100.
No. The whole point is that you upload footage, state a goal, and receive a finished edit. The manual timeline exists only for people who want to nudge the result.
It writes a per-project timeline from your content — takes, beats, pacing curve, captions and music are decided from what's in the footage, not from a preset.
Transcription, scene analysis, smart cutting, captions, audio repair, planning, rendering and the review loop form the MVP. Reframe, scoring and colour are in build; B-roll reasoning and sound design are research.
Server-side, on an FFmpeg-based render graph with background jobs. The browser only ever plays proxies and previews.
Zero editing knowledge required. The AI is the editor — you're the director.