Product and requirements
Product boundary
insonic owns orchestration, user configuration, the library catalog, processing history and the presentation of source-linked results. It integrates media tools, Cueson, model runtimes, storage and databases through explicit adapters. Local processing and hosted providers are both supported routes. Speaker-tagged audio segments across the library can form a user-elected training corpus; trained artifacts belong to their originating speaker records. Training is optional and never a prerequisite to using the library.
Users can import local files or acquire user-targeted remote sources through acquisition adapters. Acquisition adapters declare supported protocols and services; failures explain unsupported formats or authentication requirements.
The primary journey is Add media, Choose processing, Process, Explore. A user can begin without learning model servers, database syntax or pipeline internals. Advanced users can select models, endpoints, channels, storage/database backends, adapters and query parameters. The CLI defines core application operations, and the GUI wraps those operations with a parity goal; explicitly documented advanced or experimental capabilities may reach the GUI later.
Requirement map
Stable requirement IDs identify observable product behavior. The named contracts define acceptance.
| ID | Required behavior | Contract |
|---|---|---|
| R01 | Installable CLI on Windows, macOS and Linux | Technology, storage |
| R02 | CLI-first shared core; GUI installs CLI and maintains core parity, with documented advanced/experimental lag permitted | Desktop |
| R03 | Track audio and video originals, provenance and duplicates | Media |
| R04 | Acquire user-targeted sources through adapters | Pipelines |
| R05 | Correlate optimized audio derivatives with original time and channels | Media |
| R06 | Accept optional subtitles for every supported media type | Subtitles |
| R07 | Produce a correlated subtitle artifact for completed speech processing | Subtitles |
| R08 | Require Cueson as the unified subtitle layer | Subtitles |
| R09 | Configure user-defined pipelines with model and hosting overrides | Pipelines |
| R10 | Provide practical local AI defaults including faster-whisper where suitable | Technology |
| R11 | Store provider keys encrypted with usable setup and replacement | Credentials |
| R12 | Track voices, speakers, aliases and corrections | Speakers |
| R13 | Supply specialized terms and speaker names as transcription context | Speakers |
| R14 | Chunk source-backed speaker assertions through adapters into the configured graph database | Graph |
| R15 | Organize application, workspaces, media and downloaded models | Storage |
| R16 | Query assets, speakers, words and times through CLI and GUI | Graph, desktop |
| R17 | Render a graphical timeline of all library media | Graph |
| R18 | Save graph queries and render force-directed selected results | Graph |
| R19 | Offer optional AI-assisted query suggestions through chosen adapters | Graph |
| R20 | Use upstream-owned branded assets and manual scriptable updates | Releases |
| R21 | Author versioned Markdown and compile static Next.js docs | Releases |
| R22 | Ship release-matched offline docs in GUI packages | Releases |
| R23 | Use Spec Kit slices and GitHub Issues, milestones and Projects | Development |
| R24 | Keep CI turnaround below ten minutes and use human-approved squash merges with branch deletion | Development |
| R25 | Maintain Keep a Changelog and brief release highlights ending with its link | Releases |
| R26 | Retain source-timed speaker audio segments across the library and revisioned speaker corpus selections | Speakers, speaker audio |
| R27 | Optionally train speaker-associated voice models and list/fetch their artifacts through CLI contracts | Voice models, schema |
| R28 | Capture immutable raw EXIF, container, stream and other available metadata before transforms, with field-level provenance | Media, schema |
| R29 | Accept per-item origination dates, date-only values and explicit time zones through import flags and batch manifests | Media, desktop |
| R30 | Default to local filesystem storage and fully support generic S3-compatible object storage | Storage, schema |
| R31 | Default to SQLite and fully support PostgreSQL for operational state, metadata, settings and associations | Schema, storage |
| R32 | Default to LadybugDB and fully support ArcadeDB graph projections and querying | Graph, schema |
| R33 | Run configured import, transforms, attribution, jobs and elected training without mandatory per-item review | Pipelines, desktop, development |
Completion and exceptions
Speech-capable input does not require supplied subtitles. Successful processing ends with a valid correlated transcript and an exportable subtitle file. Pending, cancelled and failed processing remains visible and resumable; it is never labelled complete. Music and ambient sound classifications may skip automatic speech work. Mixed media can still contain speech, and a user can force transcription or override a mistaken classification. Neither a model's classification nor the absence of speech confidence makes a file ineligible for the library.
Keep uncertain speaker identities and uncertain calendar dates usable. The application records uncertainty explicitly instead of guessing. Direct querying and exploration work without AI query assistance enabled.
Configured routines run without repeated approvals. Only the affected operation requires additional input when a date cannot be interpreted, a destructive effect is not already authorized, or a new external processing route has not been selected. Confidence filters and attribution policies are user-configurable automation inputs; uncertain results remain labelled and correctable without mandatory review queues.
Delivery dependencies
Runtime, storage and media contracts establish the import-to-subtitle workflow on all three operating systems, with full S3-compatible storage and PostgreSQL support. Provider flexibility and speaker continuity support graph extraction, desktop exploration and optional speaker training. Model discovery and retrieval begin with CLI contracts; desktop presentation follows the shared operation model. See delivery outcomes for dependencies. Multi-user collaboration and an insonic-hosted library service are outside these application contracts.