Building a reviewable long-form video transcription pipeline
A local workflow for turning long recordings into passages I could search, review, and trace back to the source.
Field notes · Systems, tools, and lessons from the work
These are write-ups from projects that were large enough to teach me something. I cover the problem, the choices we made, who did the work, and how I checked the result. Names and operational details are left out on purpose.
A local workflow for turning long recordings into passages I could search, review, and trace back to the source.
Field Note 001 · September 2026
A transcript by itself was not enough. I needed to find a passage later, jump back to the right moment in the recording, and see what had been reviewed or corrected.
I had a collection of long videos and needed a better way to work through them. Plain transcripts gave me text, but finding something again was clumsy, and it was too easy to lose the connection to the original recording.
Every passage needed to keep its source, timestamps, processing history, and review state. If that trail disappeared, the transcript was much less useful.
I used AI agents heavily on this project. I defined the problem, set the constraints, and made the calls on privacy, review behavior, scope, and what counted as done. I also used the tool as it developed and changed direction when something looked good on paper but was awkward in practice.
AI agents wrote most of the first-pass implementation code. They did much of the test writing, repetitive data work, documentation, and debugging too. They suggested designs and worked through failures. I required working files, repeatable output, and actual test results before accepting a stage.
My part was direction, operational judgment, requirements, and acceptance. The agents did most of the implementation labor. I would not present this as a tool I coded by hand, but it also was not a hands-off prompt that produced a finished system.
We split the workflow into five stages. Each one leaves behind something that can be inspected before the next stage starts:
The review tool keeps meaningful corrections as new revisions instead of overwriting the old text. Editing a note does not disturb its classification. Changing what a passage says does, so that passage has to be reviewed again. The interface also keeps machine actions separate from human decisions.
The run processed dozens of long-form sources and produced more than 1,500 reviewable passage occurrences. The agents ran the automated and runtime checks. I used the resulting workflow, asked for changes, and decided whether each stage was ready. A successful model response was only one check:
The transcription model was the easy part to explain, but most of the work happened around it. Jobs had to resume cleanly. Failed sources needed an honest record. Timestamps had to survive every processing step. The review state could not blur an agent's suggestion into a person's approval.
Those details are what made the tool usable after the first successful run. I can return to a passage, see where it came from, see what changed, and tell whether a person has actually reviewed it.