Skip to main content

Create Spoken-Content Rough Cut

Create Spoken-Content Rough Cut builds one compact, editable sequence from transcribed spoken material. It can work from selected Project-panel clips and bins or from speech already arranged and trimmed in an active sequence.

The workflow is not limited to interviews. It is useful for presentations, tutorials, lectures, podcasts, documentary speech, video essays, talking-head material, and other edits where the spoken words define the first story pass. Workflow AI chooses one coherent path through the available transcript; local code validates every chosen passage, calculates its source range, and verifies the result created in Premiere.

New to Automation Agent?

This recipe uses Automation Agent for Adobe Premiere. Install the Python Runtime Pack, configure a Workflow AI provider, and make sure the source material already has usable Premiere transcripts.

Fastest Path: Run The Library Workflow

  1. Open Library > Examples > Editing and run Create Spoken-Content Rough Cut.
  2. In Input material, choose selected Project clips and bins or the active sequence.
  3. Describe the desired story, required material, exclusions, tone, audience, or other editorial requirements in the optional brief.
  4. Choose how much ordering freedom Workflow AI has, set a duration policy, and adjust the maximum transcript-aware handles if needed.
  5. While the setup dialog is still open, select the intended clips or bins, or activate the intended sequence. The workflow captures that scope only after you click OK.
  6. Review the measured-scope confirmation before the transcript and complete brief are sent to Workflow AI.
  7. Review the new sequence, every marker, the completion warning if present, and the omission information before using the result in a real edit.

The workflow creates a new sequence. It does not replace or trim the input sequence, and it does not alter the underlying media files.

Choose The Right Spoken-Content Workflow

These tools solve related but different editorial jobs:

Your goalUse this workflowResult
Compare every repeated scripted delivery without asking AI to choose a winner.Build Repeated-Takes Comparison SequenceA local stacked take-review sequence.
Compare semantically different alternatives while retaining several possible constructions.Create Semantic Rough-Select ComparisonA Workflow-AI-planned comparison sequence with alternatives on higher tracks.
Assemble an approved script or paper edit whose wording should drive the result.Assemble Sequence From Script (Offline)A local target-text assembly with alternatives and missing passages made visible.
Choose one coherent story path from a complete spoken-content scope.Create Spoken-Content Rough CutOne compact, sequential rough cut with review markers and an optional length goal.
Keep the current edit and only identify reusable short moments.Find Highlights And Teasers In Active SequenceScored duration markers on the active sequence; no new rough-cut assembly.

Choose the spoken-content rough cut when you want Workflow AI to make editorial selections and commit to one proposed path. Choose a comparison workflow when you want the editor, not the workflow, to decide between stacked alternatives.

Input Modes And Story Order

Selected clips and bins

Select one or more direct clips, bins, or a mixture of both in the Project panel. Eligible clips inside selected bins and all subbins are included recursively. Overlapping selections are deduplicated. Unsupported items such as sequences, multicam items, merged clips, offline sources, and material without a usable transcript are skipped and reported when other eligible material remains.

Existing Subclips are supported as input, but Premiere does not automatically inherit the master clip's completed transcript onto them. Transcribe each Subclip itself before running the workflow. Only words inside that Subclip's own source interval are eligible; the workflow does not reach back into the rest of the master clip.

With Keep existing order where it is defined, passage order is preserved inside each source clip. There is no reliable global selection order between separate clips, so Workflow AI decides how the clips relate to one another. Selecting clips in a particular click order does not establish a dependable cross-clip story order.

Material in the active sequence

Use this mode when the timeline already contains editorial decisions worth preserving. The workflow can analyze all unmuted audio tracks or one specified, one-based speech track. With Keep existing order where it is defined, the captured timeline order is global, including repeated occurrences of the same source clip.

This is the reliable way to enforce an order across several clips: arrange and trim them first, then run the workflow from the active sequence. The workflow will not recover speech outside those captured occurrence trims.

In either input mode, Allow Workflow AI to reorder all passages removes the existing-order constraint and lets the model propose a different story order. It still cannot invent words, reverse a passage, or select material outside the validated source scope.

Editorial Brief And Provider Choice

The complete usable transcript scope and the complete editorial brief are sent together for one global story analysis. The workflow does not silently split the story into independent chunks, shorten the brief, or truncate accepted transcript text. Long, detailed instructions are welcome: specify required quotes or topics, exclusions, intended audience, tone, thesis, pacing, and facts or qualifications that must survive the cut.

This global task is more demanding than sentence-by-sentence proofreading or translation. Codex CLI or Claude Code are recommended. Most current local LM Studio models are unlikely to make reliable long-form editorial decisions, even when the request technically fits. LM Studio is not categorically unsupported: use it only after the exact model and context configuration has proved suitable for comparable material.

The workflow declares a provider capability rather than locking itself to one vendor. Provider limits still apply. Before sending anything, it shows the eligible scope and transcript/brief size; if the complete request cannot be transported or processed, reduce the Premiere scope or use a more capable provider. The workflow does not create a sequence from a truncated substitute.

Privacy and cost

Transcript text and the editorial brief can contain sensitive production or client information. They are processed according to the configured Workflow AI provider and may involve network use, provider token costs, and several minutes of analysis. Confirm that the chosen provider is approved for the material.

Duration Goals Are Measured, Not Promised

Choose one of three policies:

  • Target with tolerance asks for a duration inside the target plus/minus the allowed variance.
  • Maximum only asks for a non-empty result no longer than the maximum.
  • No fixed target asks for the shortest coherent cut that satisfies the editorial requirements without enforcing a numeric duration.

Timing is calculated locally from the exact materialized word ranges. When a valid first plan misses a fixed goal, the workflow may ask Workflow AI once for a complete corrected plan. If the safest usable result still misses, it creates that sequence with LENGTH MISMATCH in the name, adds a warning marker, and reports the accepted interval, actual duration, and measured miss. It does not hide a useful result merely because its duration is imperfect.

An unsafe plan, unknown source ID, duplicated source use, invalid range, or other structural failure is different: those errors stop the workflow before sequence creation.

Handles Avoid Neighboring Words Whenever The Frame Grid Allows It

The before/after values are maximum padding, not guaranteed additions. Each handle is shortened or omitted when it would cross:

  • the immediately preceding or following unselected transcript word,
  • a source-media boundary,
  • an active-sequence occurrence trim, or
  • a frame boundary that cannot represent the full requested padding safely.

The workflow preserves the selected speech and never includes an adjacent unselected word merely to reach the requested handle length. One unavoidable case is different: a transcript word boundary can fall inside a Premiere video frame. If moving inward would trim selected speech, the workflow keeps the smallest outward frame instead of discarding the whole rough cut. That frame can contain a few milliseconds of the neighboring word. The Sequence receives a warning marker, and the report states the exact extension and known neighboring- word overlap.

What The Workflow Creates

After the plan and all source ranges pass local validation, the workflow:

  • creates a matching sequence beside the source used for sequence settings. In selected-clips-and-bins mode, the first planned placement with video supplies those settings, falling back to the first planned placement when the complete plan is audio-only. In active-sequence mode, the captured sequence supplies the settings;
  • places the validated backing-source ranges sequentially on V1/A1 as available;
  • creates review markers for sections, uncertainty, brief conflicts, and a possible duration mismatch; if Premiere rejects an individual marker, it retains the clip sequence and reports the missing marker count;
  • verifies requested audio/video lanes, placement counts, and observable source-range restoration before declaring the clip edit complete; and
  • optionally retains a detailed Markdown report and troubleshooting files under ::HOME::/automation-agent-spoken-content-rough-cut.

The durable report includes the available request-size, coverage, duration, attempt, warning, and result information.

Clean Reconstruction Limits

Active-sequence mode uses the timeline as a speech-selection and ordering source. It reconstructs the chosen backing-source ranges into a clean new sequence; it does not clone or refine the existing timeline construction.

The new sequence does not preserve:

  • clip or track effects and automation,
  • transitions,
  • graphics and lower thirds,
  • music, captions, or B-roll selected through visual meaning,
  • multicamera switching,
  • nested-sequence internals, or
  • other non-speech timeline construction.

Muted or disabled material and occurrences without usable transcripts are reported rather than treated as spoken candidates. Nested sequences, multicam clips, and merged clips are skipped. An in-scope active-sequence occurrence with reverse playback or a non-1x speed instead blocks the complete run before Sequence creation; isolate normal-forward 1x speech on the chosen track, or mute/disable the affected occurrence and rerun. Distinct eligible speech occurrences that overlap in timeline time are ambiguous in the all-unmuted- tracks mode and also block the run; isolate the intended speech track and run again.

Silent or visually driven meaning is outside this transcript-only workflow. If a title card, reaction shot, demonstration, screen recording, music beat, or silent reveal changes the story, restore or edit it manually after reviewing the rough cut, or use a live agent workflow that deliberately inspects visual evidence.

Failure And Review Behavior

Before sequence creation, invalid setup, insufficient transcript coverage, provider failures, and rejected plans leave Premiere unchanged.

Once mutation begins, the new sequence is staged with INCOMPLETE in its name. Each requested placement and marker is verified before the next success state. Frame-grid boundary extensions and a length status changed by exact frame alignment are usable rough-cut imperfections: the workflow continues, adds warning markers, records them in the report and completion message, and removes the incomplete label after final verification. A marker-creation failure is also non-fatal because it does not invalidate the placed clips; the completion message and console state exactly how many markers are missing. If Premiere instead reports a restoration, placement, or read-back mismatch, the workflow stops and retains the clearly labeled incomplete sequence plus diagnostics rather than pretending a potentially corrupted or missing placement is complete. If an active-sequence trim or source-media boundary cuts through a selected spoken word, the error identifies the blocking boundary and its measured time, and suggests extending the Timeline occurrence trim or using the source clip.

The verified Sequence is finalized before the optional durable report is copied. A filesystem error while copying that report can therefore be reported without turning the already verified clip edit back into an incomplete output.

Even a technically complete result remains an editorial proposal. Review:

  • every cut and transition in meaning,
  • omitted material and repeated source use,
  • qualifications, uncertainty, and factual dependencies,
  • transcript timing and shortened handles,
  • every review marker, and
  • any LENGTH MISMATCH or INCOMPLETE output before further use.