insonic · v0.0.0

insonic v0.0.0

insonic organizes audio and video, correlates subtitles, tracks speakers and explores what was said and when through a CLI and desktop wrapper. Replaceable local and hosted tools operate through explicit adapters while the user controls the library, configuration and processing history. Source-linked speaker audio can optionally feed trained models associated with those speakers.

v0.0.0 is a specification baseline. Repository tooling and the documentation site exist; application commands and installable CLI/GUI packages described here are contracts for implementation.

Read in this order

  1. Product and requirements states the outcome and required behavior.
  2. Architecture explains authority, processing and recovery boundaries.
  3. Technology decisions compares implementation choices and dependency qualification.
  4. Delivery outcomes explains capability dependencies.

Focused contracts cover import dates and metadata, media, pipelines, subtitles, speakers, speaker audio and models, queries, storage, catalog structure, JSON contracts, credentials, desktop behavior and development. The glossary defines technical vocabulary with learning and local usage links. The changelog records changes.

The guiding outcome

A user adds media with original metadata and optional recording dates/subtitles, selects an understandable preset and receives a searchable library. Results link to the original media position and distinguish acoustic voices, named speakers, transcript wording and extracted assertions. The CLI exposes core operations and machine-readable results; the GUI wraps the shared runtime and provides a media timeline and saved graph views. Optional speaker-model training uses an immutable corpus snapshot without interrupting ordinary library work.

Diagram source
flowchart TB
  Import[Audio or video with optional subtitles] --> Library[Tracked original and derived assets]
  Library --> Pipeline[User-selected local or hosted pipeline]
  Pipeline --> Cues[Correlated Cueson transcript and speaker evidence]
  Cues --> Query[Search and source-linked graph queries]
  Cues --> Corpus[Speaker-linked audio corpus]
  Corpus --> Training[Optional elected model training]
  Training --> Models[Versioned speaker models and CLI retrieval]
  Query --> CLI[CLI results]
  Query --> GUI[Desktop timeline and saved graph views]