Field notes · Systems, tools, and lessons from the work

Notes from the work.

These are write-ups from projects that were large enough to teach me something. I cover the problem, the choices we made, who did the work, and how I checked the result. Names and operational details are left out on purpose.

Latest note

Field Note 001 · September 2026

Building a reviewable long-form video transcription pipeline

A transcript by itself was not enough. I needed to find a passage later, jump back to the right moment in the recording, and see what had been reviewed or corrected.

The problem

I had a collection of long videos and needed a better way to work through them. Plain transcripts gave me text, but finding something again was clumsy, and it was too easy to lose the connection to the original recording.

Every passage needed to keep its source, timestamps, processing history, and review state. If that trail disappeared, the transcript was much less useful.

How the work was split

I used AI agents heavily on this project. I defined the problem, set the constraints, and made the calls on privacy, review behavior, scope, and what counted as done. I also used the tool as it developed and changed direction when something looked good on paper but was awkward in practice.

AI agents wrote most of the first-pass implementation code. They did much of the test writing, repetitive data work, documentation, and debugging too. They suggested designs and worked through failures. I required working files, repeatable output, and actual test results before accepting a stage.

My part was direction, operational judgment, requirements, and acceptance. The agents did most of the implementation labor. I would not present this as a tool I coded by hand, but it also was not a hands-off prompt that produced a finished system.

The constraints

  • Recordings varied in length, audio quality, and availability.
  • Processing needed to run locally with GPU acceleration rather than sending the source collection to a hosted transcription service.
  • Interrupted jobs had to resume without repeating completed work or silently losing state.
  • Machine-generated candidates could assist a reviewer, but could not be represented as human-approved conclusions.
  • The review interface needed to remain useful from a separate computer on the trusted network while the processing workstation stayed headless.

What we built

We split the workflow into five stages. Each one leaves behind something that can be inspected before the next stage starts:

  1. Ingest and inventory. Record the source manifest, availability, and failure state before expensive processing begins.
  2. Extract and transcribe. Normalize audio and run local GPU-assisted speech recognition in restartable batches.
  3. Segment with provenance. Split transcripts into timestamped passages without discarding their source identifiers.
  4. Promote candidates. Apply controlled search and classification steps to create a smaller review queue while preserving the path back to the full transcript.
  5. Review in context. Present the passage, timestamps, source metadata, notes, classifications, and revision history in a lightweight browser interface.

The review tool keeps meaningful corrections as new revisions instead of overwriting the old text. Editing a note does not disturb its classification. Changing what a passage says does, so that passage has to be reviewed again. The interface also keeps machine actions separate from human decisions.

Verification

The run processed dozens of long-form sources and produced more than 1,500 reviewable passage occurrences. The agents ran the automated and runtime checks. I used the resulting workflow, asked for changes, and decided whether each stage was ready. A successful model response was only one check:

  • Source manifests preserve unavailable and failed items rather than quietly omitting them.
  • Promoted passages retain timestamps and stable links back to their source records.
  • Write-ahead recovery protects review work from an interrupted save.
  • Host and origin validation limit how the browser interface can be reached.
  • Automated tests cover revision behavior, stale-classification invalidation, navigation, and audit history.

What I learned

The transcription model was the easy part to explain, but most of the work happened around it. Jobs had to resume cleanly. Failed sources needed an honest record. Timestamps had to survive every processing step. The review state could not blur an agent's suggestion into a person's approval.

Those details are what made the tool usable after the first successful run. I can return to a passage, see where it came from, see what changed, and tell whether a person has actually reviewed it.