Articles / Viewpoints and methods
7 minFor tool users

Narration and Scene Timing: A Manual Audio-First Worksheet

Plan scene cuts from audio durations, distinguish separate chapter files from one continuous track, and review transitions. Includes a copyable blank worksheet and a fictional worked example.

Aaron HuangSystems, product and AI practice

Narration and Scene Timing: A Manual Audio-First Worksheet

If the narration is still explaining one chapter while the video has moved to the next, build the scene schedule around the audio: measure each chapter, record where pauses belong, then set the visual cuts. Script length can help estimate scope, but it cannot replace the duration of the audio you will actually use.

This guide describes a manual scheduling method, with a copyable blank table and a three-scene worked example in the article. No narration was generated, no editing software was operated, and no export was tested for this article. Every number in the example is fictional teaching data.

Is this a chapter mismatch or a different audio problem?

The task here is to show the chapter that the narration is discussing. If the voice starts explaining step two while the screen still shows step one, chapter boundaries and transition times give you something concrete to compare.

This need appears in html-video Issue #36 and Issue #66. The same user first asked how to align variable-length narration with visual cuts, then asked about automatic alignment with article chapters. These are two reports from one user. They do not establish how widespread the problem is or prove that the requested feature has been implemented.

The html-video README’s Soundtrack section describes narration, background music, and mixing audio into the export. That description does not establish automatic chapter-by-chapter synchronization. Generating narration still leaves chapter timing to be checked.

If your problem is a speaker’s lip movements, Bluetooth or device playback latency, or sample/clock drift that increases during playback, identify that cause first. This chapter worksheet does not provide fixes for those problems.

Measure an audio file you have permission to use

Start with a copy of narration you have permission to use. Fix its filename and version before measuring it. Regenerating narration, changing its speed, removing a sentence, or trimming silence can move later cut points.

If you already use Audacity, select the file or a chapter range and read its start, end, or length in the Selection Toolbar. The manual lists displays such as Start and End and Start and Length. Record times consistently in seconds and milliseconds: 1 minute 30 seconds is not 1.30 seconds. The selection gives you a duration; you still need to listen to decide which sentences belong to each chapter. Audacity Selection Toolbar manual

If ffprobe is already available in your environment, the following command can request the container duration of an audio file. This is a documentation-based command example, not a command executed for this article. Replace the filename with your own copy:

ffprobe -v error -show_entries format=duration -of default=noprint_wrappers=1:nokey=1 "chapter-01.wav"

The result records whole-file duration. It does not identify chapters or determine when speech ends. If no usable value is returned, mark the duration as awaiting measurement rather than entering zero or estimating from word count. The file may include leading or trailing silence. If you trim it, measure the version you actually use again. ffprobe documentation: show_entries and output formats

One file per chapter: add each extra pause once

For this example, chapter files play in sequence without overlap, starting at video time zero, with direct visual cuts. “Audio duration” includes silence already in the file. “Added pause after chapter” means only extra time inserted after it. Do not count a pause already inside the file twice.

Use these calculations:

  • The first scene starts at 0 seconds
  • Scene end = scene start + chapter audio duration + added pause after the chapter
  • The next scene starts when the previous scene ends

The following values were created for this example. They do not correspond to real audio files or playback results:

Scene Audio duration (s) Added pause (s) Scene start (s) Scene end (s) Visual plan
1 Problem 8.4 0.6 0.0 9.0 Show the problem through the end of the pause
2 Method 12.2 0.8 9.0 22.0 Show the method and steps
3 Conclusion 7.5 0.5 22.0 30.0 Show the conclusion and hold through the final pause

This schedule keeps the preceding visual on screen during each pause. It plans cuts at 9 and 22 seconds and ends at 30 seconds. The audio totals 28.1 seconds; added pauses total 1.9 seconds. Their sum is 30 seconds. That checks the arithmetic, not whether the visuals change at the right point in the explanation.

Start with direct cuts to make the schedule clear. If you later add fades, record each transition’s start and end as well. Do not automatically add an overlapping visual transition to the total runtime as extra time. Follow the timeline in the tool you actually use.

One continuous track: keep the original timestamps

If the narration is already one continuous track containing all chapter pauses, mark the chapters on that track. Do not add the chapter durations and existing pauses again to its original start timestamps.

Using the same fictional example, suppose the track starts at video time zero. The spoken chapter ranges are 0–8.4, 9–21.2, and 22–29.5 seconds. The pauses between chapters and at the end already exist in the 30-second track. Scenes can still run from 0–9, 9–22, and 22–30 seconds. Enter 0 for added pauses and describe the existing pauses in the notes.

The second scene starts directly at the track’s 9-second mark. Adding the preceding 9 seconds again would incorrectly move it to 18 seconds. If the whole track starts 2 seconds into the video, add that same 2-second offset to every original cut point. If you remove a section or insert an actual pause in the middle, the timeline has changed: mark the affected chapters again.

Fill the worksheet so you can locate the mismatch

The blank table below uses seconds throughout. Copy it and fill it in manually; it does not recalculate automatically. Leave missing values blank and mark them as awaiting measurement; do not use zero to mean unknown.

Field What to record
Chapter and audio mode The chapter, and whether you use one file per chapter or one continuous track
Audio file and version / source start and end Enough information to find the same audio again; source timestamps are required for a continuous track
Measured duration The duration of the separate file used, or the selected chapter’s source end minus source start
Added pause after chapter Extra time not already included in the file or continuous track
Scene start and end / visual Planned positions on the video timeline and the content to show
Actual transition start and end Recorded when reviewing the output; use the same time for a direct cut, and leave blank until reviewed
Revision status Awaiting measurement, awaiting review, needs adjustment, or reviewed, with the specific issue in notes

Leave actual transitions blank until you review the output. The calculated scene boundaries in the example above are not measured cut times from an exported video.

Copyable blank table

Use one row per chapter. Enter both start and end values in each range field; all times are in seconds. Leave unknown numeric values blank and write “awaiting measurement” in the status field.

Chapter / audio mode Audio file and version Source start / end Measured duration Added pause after chapter Scene start / end Visual Actual transition start / end Revision status / notes

Review both sides of every transition after export

Do not stop at “the audio and video are both 30 seconds.” Their total lengths can match while the second chapter appears too early or too late.

Watch and listen around each cut, starting with perhaps 1–2 seconds on either side. This is an initial review window; extend it for longer sentences. Check that the previous chapter’s closing sentence still has the right visual, and that the next chapter’s visual is ready when its opening sentence begins. Record the actual transition times. Then review the whole video, including its opening, ending, and important visuals within each chapter.

If the screen changes before the sentence ends, compare the chapter’s audio duration, any missing pause, and whether the tool applied the planned timing to the output. After changing visuals, export and review again. After changing narration, start again with audio measurement and recalculate the affected later scenes.

The worksheet lets you identify which chapter, second, and visual do not line up. Mark that version as checked only after listening, watching, and recording the corrections.