Files
Miguel Palhas 6cd2902341
ci / web (push) Successful in 1m0s
e2e / e2e (push) Failing after 3m17s
ci / rust (push) Successful in 4m0s
ci / image (push) Successful in 3m56s
feat(subs): translate from an image-only release
A release whose only subtitle is a PGS or VobSub track was stuck both
ways: the track satisfied English so no English SRT was ever fetched,
and bitmaps can never feed a translator. §15 is amended to separate
satisfying viewing from providing a translation source, and to carve a
fetch made to obtain a source out of the no-upgrade rule.

Timings come from the disc rather than from alass guessing at the audio.
At import, each non-forced image track's packet timestamps are paired
show-to-clear into a cue skeleton and stored; the fetched source is then
aligned against that skeleton, and the translation made from it skips
the post-translation pass, which could only move disc-exact timings off.

Pairing is validated before it is trusted — even packet count, plausible
durations, sane density for the runtime — because PGS allows several
composition segments per subtitle and a slipped pairing is quietly half
a second out. A track that fails validation gets no skeleton and falls
back to aligning against the video, as does any file imported before
this: there is no backfill.

Closes #268

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-30 11:53:26 +01:00

50 KiB

arr — design

One service that replaces Radarr and Sonarr, and later Bazarr, for a single household. Rust workspace, API-first, SQLite, no authentication layer.

This document is the contract. Issues reference it by section rather than restating it.


1. Why

Five .NET services idle at 600-900 MB to do a job with one loop in it: decide what you want, find it, fetch it, put it somewhere Jellyfin can read. Each has a configuration surface an order of magnitude larger than the subset actually used, and none of them share a model, so the same show is configured three times in three dialects.

The replacement targets 20-40 MB resident, one binary, one database, one policy language.

This is not a general-purpose Radarr replacement. Every decision below is allowed to be narrow.

2. Non-goals

Explicitly out of scope, permanently unless stated:

  • Authentication. The perimeter is a VPN plus Authelia at the proxy. The service binds without auth, same trust model as the rest of the stack.
  • Library migration or filesystem scan. The service knows only what it put on disk. Adopting the pre-existing library, if ever wanted, is a one-off script against both APIs, not a feature.
  • Automatic quality upgrades. See §5.4.
  • Multiple quality profiles. Policy attaches to a root folder. See §5.1.
  • Trakt. Jellyfin and Stremio already write watch state there and nothing in this pipeline reads it.
  • Calendar view, actor/director following, per-indexer UIs. Prowlarr owns indexer configuration and keeps doing so.
  • Books and music. shelfmark and calibre-web-automated already cover books. If they ever land here they reuse maybe three of the ten crates, so nothing is generalised in advance. See §11.

3. System context

                    ┌──────────────┐
        TMDB ──────▶│              │
                    │     arr      │────▶ qBittorrent WebUI  qbittorrent.n62.casa
    Prowlarr ──────▶│              │
   (Torznab)        │   (this)     │────▶ ffprobe            local subprocess
                    │              │
   Jellyseerr ─────▶│              │────▶ Jellyfin refresh   10.6.10.18:8096
   (Radarr-compat)  └──────┬───────┘
                           │              ntfy                per-person topics
                           ▼
                    SQLite + media tree on /mnt/media

Everything already exists except arr. Prowlarr keeps owning tracker auth, Cloudflare bypass via FlareSolverr, rate limiting and the Cardigann definitions — replacing it buys nothing.

qBittorrent is reached at qbittorrent.n62.casa, WebUI API v2, download dir /mnt/media/qbittorrent/complete. Its WebUI requires a login, so arr carries credentials — the one upstream that does.

4. Domain model

Movies and series are separate aggregates. They share machinery, not a table.

Root
  id, kind (movie|tv), audience (main|kids), path, policy_id

Policy
  id, name
  required_audio          rule set, see §5.2
  dub_blacklist           [pt-BR]
  hdr_rules               see §5.3
  size_bands              per resolution: floor, target, penalty curve,
                          all per episode (§5.5)
  resolution_pref         [2160p, 1080p]
  source_weights          small tiebreaker

Movie
  id, tmdb_id, title, year, original_language, root_id
  wanted (bool), overrides (json), state, blocked (bool)
  search_attempts, last_searched_at

Series
  id, tmdb_id, title, year, original_language, root_id
  auto_track (bool), overrides (json), blocked (bool)

Season
  id, series_id, number, tracked (bool)

Episode
  id, season_id, number, title, air_date
  wanted (bool), state
  search_attempts, last_searched_at

MediaFile
  id, owner_kind (movie|episode), owner_id, path, size
  probed (json: resolution, source, hdr, audio_tracks, sub_tracks)
  waiver (json|null)   which rule was relaxed to allow this import

Release
  id, indexer_id, guid, name, size, seeders, publish_date, download_url
  parsed (json)        what the name claims
  score, verdict       eligible | waived | rejected(rule)

Grab
  id, release_id, target_kind, target_id, infohash
  state, grabbed_at, imported_at

Blacklist
  infohash, normalised_name, reason, created_at

Owner
  id, name, ntfy_topic

TitleOwner
  title_kind, title_id, owner_id      many-to-many

4.1 Intent lives at the leaf

Sonarr's monitored conflates "do I want this show" with "is anything happening right now". It is user-maintained, so it drifts and stops meaning anything.

Here, intent is Episode.wanted and Movie.wanted only. Series.auto_track is not intent — it is a rule one level up: when metadata reveals a new season, that season becomes tracked. The rule applies from the second refresh onward: seasons revealed by a series' first refresh — the whole back catalogue present at add time — are never tracked by it. Only seasons that appear after the series was added are covered; back-catalogue seasons are picked by hand. An untracked series where you manually marked S02 needs no special case: three wanted episodes, nothing else.

Season.tracked is also a rule, not intent. Turning tracking on marks every already-revealed episode in the season wanted, and keeps marking episodes wanted as metadata reveals them. Turning it off clears wanted on every episode of that season: toggling the flag is itself the explicit act, and the product offers no other way to withdraw a season's worth of intent. Without that, untracking leaves a series with every season off and every episode still wanted, pinned at incomplete with no way back. Marking individual episodes wanted on a season that was never tracked is unaffected by any of this.

Both rules skip season 0. Specials are dozens of undated shorts and recaps that indexers do not carry, so auto_track never marks season 0 tracked. See §4.2 for how season 0 stays out of derived status.

4.2 Status is derived, never stored as intent

Recomputed on metadata refresh and on file change:

Status Condition
airing an episode aired or airs within ±14 days
incomplete wanted episodes missing, nothing currently airing
waiting season finished, next unannounced or future-dated
complete everything wanted is on disk
ended series finished upstream and complete

Default list view shows airing and incomplete. Everything else collapses behind one toggle. A one-off season grab therefore disappears from the default view by itself once satisfied, and a tracked show reappears by itself when a new season is announced. Nothing to remember to flip.

Status chips carry state colour, not a single informational hue: green (--signal-ok) for on disk or complete, amber (--signal-warn) for wanted but not yet found, violet (--status-airing) for downloading, neutral for untracked. The word stays present in every chip, so state survives with colour removed.

One state is exempt: a file being on disk is drawn as a check glyph rather than the word available, with the word on the chip's aria-label. A drawn check is a shape, not only a hue, so it survives the greyscale read the rule exists to protect. The exemption is that narrow — every other state, and every derived-status chip above, keeps its word, because no unambiguous glyph stands in for incomplete or waiting.

Season 0 is invisible to all of it: derived status ignores season 0 episodes entirely, so a manually wanted special cannot pin a series at incomplete or hold back ended. The accepted consequence is that specials are visible and grabbable in the series detail view, but never affect status.

4.3 People are tags, not paths

Owner is a many-to-many tag on titles. It drives UI filtering, the default list, and notification routing. It never appears in a path. A file lives in one place; audiences overlap.

5. Policy engine

The core of the app and the part worth testing hardest. Lives in arr-core, pure, no IO.

5.1 Policy attaches to a root

Two roots per media kind (main, kids), each with one policy. Jellyseerr's "pick a quality profile per request" model is satisfied by exposing exactly one fake profile per root and ignoring whatever it sends.

Per-title overrides relax the root policy for one title. Same mechanism in both directions — only_4k tightens, allow_english_audio loosens.

Stored verdicts follow the effective policy. A release's verdict is stamped by the search that found it, and both §9.3's deck and the daemon's manual-grab gate read that stored column. So the four operator actions that change a title's effective policy — editing its overrides, moving it to a root with a different policy, pointing a root at a different policy, and editing the contents of a policy some root points at — re-derive the stored verdicts of everything they touch, inline in the same request. The operator is never left reading a verdict computed under a policy that no longer applies.

The fourth is the widest: a repoint moves one library, a policy edit moves every library sharing the policy. Inline still holds there. Re-evaluation is pure and in-memory, a row whose verdict does not move is not rewritten, and the ceiling is the whole database rather than something that grows with the number of roots — roots partition titles, and a title has exactly one root. Measured over 2000 titles and 10 000 stored releases split across two roots sharing one policy: 0.36 s when the edit moves no verdict, 2.7 s in the pathological case where it moves all 10 000. A rename changes no rule and re-derives nothing.

5.2 Language

Requires a concept the release name does not carry: the original language of the title, from TMDB, distinct from a track's language.

reject audio track T if  T.lang ∈ dub_blacklist  and  T.lang ≠ title.original_language

dub_blacklist = [pt-BR] globally. That single expression covers both roots:

  • main — require a track matching original_language. Others may ride along. A Brazilian film has original_language = pt-BR, so its own soundtrack passes; a pt-BR dub of an English film does not.
  • kids — require pt-PT, or original_language when that is already Portuguese. pt-BR never satisfies the requirement, because the child does not read and a Brazilian dub is not what he should be watching.
  • subtitles — no blacklist. pt-BR subtitles are always fine.

The pt-PT / pt-BR detection problem. ISO-639-2 has one code, por, for both. ffprobe reports por either way. Three signals in order:

  1. Release name markers — Brazilian releases self-identify loudly: PT-BR, Dublado, Nacional, Dual Áudio.
  2. ffprobe stream title and handler_name strings, not the language code — often literally Portuguese (Brazil).
  3. Container-level BCP-47 tags where present (Matroska can carry pt-BR).

When none resolve it, tag the track por-unverified and surface it. Do not guess. Guessing toward kids gives a child a Brazilian dub; guessing toward main rejects a Brazilian film's own audio.

Sourcing consequence for kids. European Portuguese dubs live almost entirely in streaming WEB-DLs (Disney+, Netflix, Max) and essentially never in BluRay encodes or remuxes. That root biases hard toward WEB-DL sources, and unlike audio in general it is often pre-grab visible, since multi-audio releases advertise themselves (MULTi, DUAL, explicit language lists).

When nothing qualifies, the title sits in a visible no-PT-source queue rather than being grabbed in English. That queue is acted on manually, sometimes for months. A per-title allow_english_audio override empties one item from it with one click, and the resulting file is recorded with a waiver so the UI shows it as "English, no dub" rather than a clean match.

5.3 HDR and Dolby Vision

The naive rule "blacklist Dolby Vision" is wrong and would discard a large share of good 4K releases.

Profile Base layer Behaviour without DV support Verdict
5 none (ICtCp) green/purple, unwatchable reject
7 dual-layer BL+EL depends on player, unreliable reject
8.1 HDR10 plays as HDR10 accept

The existing library is already almost entirely Profile 8.1 with no Profile 5, and Jellyfin tonemaps it correctly on the current GPU path.

Profile is knowable only from ffprobe, never from a release name. A DV chip in search results means "the name claims DV" and its absence means nothing.

5.4 Quality, and why there is no upgrade loop

Default: take 4K if it exists now, otherwise 1080p. Per-title only_4k rejects 1080p outright and keeps searching indefinitely — the escape hatch for the handful of titles worth waiting on.

No grab delay. Deliberately rejected. Radarr's delay profile exists to stop first-match-wins from always yielding 1080p; here the only_4k flag solves the same problem upfront and without making every grab slower.

No automatic upgrade loop. An automatic upgrade re-downloads 40-80 GB and restarts a seeding obligation for something already watched. needs_upgrade is a marker in the UI and a manual "search again" button, not a background job.

5.5 Scoring: distance from a target size

Source tier is not the dominant term. A remux should win only when nothing smaller exists.

Per resolution, a floor, a target and a growing penalty above target. For 4K: floor ~8 GB, target ~22 GB. A 20 GB WEB-DL sits at target and wins; a 60 GB remux scores badly but stays eligible, so it is picked when it is the only option. The floor matters — unbounded "smaller is better" selects a 3 GB 4K encode that looks like mud.

A band describes one episode. For a movie the question never arises — one release is one film. For TV a release may carry a single episode, several, or a whole season, and the shipped TV numbers are episode numbers: a 1080p target of 2 GiB is what one episode should weigh, not a ten-episode pack. So a release is measured by its size divided by the number of episodes it covers, and that per-episode figure is what both the target penalty and the floor are compared against. A 20 GiB pack of ten episodes is scored as 2 GiB and sits at target, not ten times above it.

The floor takes the same figure. It is a hard reject rather than a penalty, so comparing a pack's total against an episode-sized floor lets every pack through untested — wrong in the opposite direction from the target penalty. The floor exists to keep a 3 GB 4K encode out, and a pack of ten such encodes has to fail it just as plainly.

An unknown episode count is one episode. A release matched before its season's episodes have been revealed by a metadata refresh has nothing to divide by, and the rule does not guess a number: the divisor is 1 and the release is measured by its full size. That is the behaviour today, and it makes a pack look oversized rather than undersized — it fails toward rejecting a good pack rather than grabbing a bad one, and the next search after a refresh has the real count.

A band describes a rate, not a fixed size per episode. The shipped values are read against a reference runtime of 45 minutes: a 1080p target of 2 GiB means 2 GiB per 45 minutes of episode. Before the per-episode figure is compared against them, a band's floor and target are both scaled by runtime / 45, where runtime is the series' minutes-per-episode from TMDB metadata. For a typical drama around 45 minutes the shipped numbers keep exactly their current meaning; a 22-minute show is judged against roughly half the floor and half the target instead of being rejected for weighing half of what an hour of video weighs. The shipped band values themselves do not change — the reference runtime is chosen so they do not have to.

A missing or zero runtime is the reference runtime. When metadata carries no per-episode runtime, or carries zero, the scale factor is 1 and the band applies unscaled — exactly today's behaviour. The rule does not guess a duration, for the same reason the unknown episode count does not guess a number.

Movies are not scaled. This applies to episodes only. A movie's bands are already tuned against feature length, so its floor and target keep their current meaning regardless of the movie's own runtime.

Source tier (Remux > BluRay > WEB-DL > WEBRip > HDTV) survives as a small tiebreaker. Seeders are log-scaled and small: enough to complete, past that it does not matter. Telesync, CAM and screener are hard filters, not low scores.

How resolutions rank against each other. Size is scored against the band for the release's own resolution, so on its own it says nothing across resolutions: an at-target 4K and an at-target 1080p both score the top of the size term, and a 4K a couple of gigabytes over target loses to a 1080p that is merely on target. That is wrong — resolution_pref is an ordered list, and the order is a preference, not just an eligibility filter.

So each step up resolution_pref is worth a fixed number of points. The last entry is worth nothing and every earlier one a step more. A resolution the list does not carry scores nothing rather than being penalised, the same as an unclaimed resolution: no ranking is no opinion.

The step is sized against the size term, not chosen in isolation. With the seeded 4K band — target 22 GB, 60 points per gigabyte over — a step of 300 is five gigabytes of overshoot: a 4K up to about 27 GB beats an at-target 1080p, and a bloated 40 GB 4K does not. Below target the same arithmetic asks a 4K to be within about four gigabytes of its target to win, which is what keeps an 8 GB 4K that looks like mud from beating a good 1080p.

Exact numbers are policy rows, tuned by hand. The model is the decision.

5.6 Two phases of truth

Release names lie or omit. Files do not.

  • Pre-grab — parse the name. It is all there is. Use it for hard filters (source type, resolution, explicit language markers) and for scoring.
  • Post-downloadffprobe before import. Real audio languages, real HDR format and DV profile, real codec, real duration.

Store both, parsed on the release and probed on the file. The gap between them is also data: it tells you which release groups mislabel.

5.7 Hard fail versus soft fail

A policy violation found by ffprobe is not one thing.

  • Hard fail — useless. No original-language audio track at all, DV Profile 5, wrong title, corrupt. Do not import, blacklist the release, grab the next candidate.
  • Soft fail — watchable but not what was asked. 1080p while only_4k was set, PT audio missing on a kids title. Import it, record a waiver, surface it.

Neither deletes the torrent. See §7.3.

Two hard failures make a decision, and only for 30 days. A movie, an episode or a season enters the needs-a-decision queue (§9.5) when two grabs against different releases hard-failed on it, and both of those failures happened within the last 30 days. One bad torrent is not a decision — a release that hard-failed is blacklisted (§6.3) and the next candidate is grabbed, which is the system working.

The window runs from the failure, not the grab. The two are usually minutes apart, but a torrent can sit stalling on a dead swarm for five weeks before ffprobe finally condemns it — and that failure is fresh evidence the target is broken now, not history. Measured from the grab it would be born outside the window and a genuinely broken target could never surface. So grabs records failed_at alongside grabbed_at, and the window reads it.

The window is what lets the queue be emptied. Nothing clears a grabs row, so without it the queue only ever grows and the one season that wants attention sits behind eight that were dealt with months ago. It is the queue's version of §6.2's "it never gives up entirely, it goes quiet": a target the operator has dealt with stops producing failures and drops out once the last one ages past 30 days, while a target that is still broken keeps producing them — the pack guard retries at worst weekly (§6.2) — and stays queued for exactly as long as it is genuinely broken. Nothing is dismissed by hand and no acknowledgement state is stored, so there is no second thing to keep correct.

Only a target still waiting for a file is queued. The count and the window already express one rule — bounded attention — and this is another face of it: a failure history queues a target only while that target still has a gap to fill. A movie or an episode is queued while it is wanted and not available. A season holds no intent of its own (§4.1), so it is queued while at least one of its episodes is still wanted and not available. A season pack that hard-failed twice, fell back to per-episode grabbing exactly as §6.2 says it should, and was then fully acquired leaves at once rather than waiting out the 30 days — that is the system working, not a decision. A target that is still broken keeps producing failures and stays.

The same rule applies on all three lanes and in both readers. GET /api/queues/attention (§9.3) and the ntfy notification (§9.5) are two views of one queue; filtering differently tells the operator two different stories on two channels.

6. Sourcing

6.1 Prowlarr, per-indexer Torznab

GET /api/v1/indexer enumerates. Each indexer is then addressed at its own Torznab endpoint:

GET /{indexerId}/api?apikey=…&t=caps
GET /{indexerId}/api?apikey=…&t=search&q=…
GET /{indexerId}/api?apikey=…&t=movie&imdbid=…
GET /{indexerId}/api?apikey=…&t=tvsearch&tvdbid=…&season=&ep=

Prowlarr's aggregate /api/v1/search is deliberately not used: Torznab is the stable standard rather than an internal UI API, it gives real RSS semantics, and per-indexer addressing keeps indexer identity in hand — which the per-tracker seeding rules need anyway.

t=caps per indexer records which search modes each tracker supports. Not all do ID-based search; some are text-only.

6.2 Three triggers, three cost profiles

Trigger Cost Cadence
RSS sync one call per indexer, independent of library size ~10 min, never backs off
Targeted search one call per wanted item per attempt exponential backoff
Manual one call, user-initiated on demand

RSS is an empty-query Torznab call whose results are matched against the wanted list locally. Because its cost does not scale with the wanted list, every wanted item is matched against every RSS result forever — with one guard: the RSS lane skips a season pack for a season that already has episodes on disk, the same guard §14 applies to re-grabs. Both lanes agree, so a season does not behave differently depending on which lane sees a release first. A pack for a season with nothing on disk stays eligible. When a single episode cannot be found on its own, the escape hatch is the season release deck, not a lane exception.

Targeted search backs off 1h → 6h → 1d → 3d, capped at 7d, reset when the title's metadata changes. It never gives up entirely, it goes quiet.

The ladder runs from the failure, not the grab. A failed season-pack grab quiets the pack lane on that same curve, and the rung is measured from the moment the grab entered failedgrabs.failed_at, the column §5.7's attention window reads — not from when it was sent. The two are usually minutes apart, but a torrent can stall on a dead swarm for five weeks before ffprobe condemns it at import. Measured from the grab, the whole ladder has already elapsed by the time the failure lands, so the lane retries the source that just failed at once, which is the one thing the backoff exists to prevent. The ladder's job is to stay off a source that has recently failed, and "recently" can only mean recently failed.

One anchor covers both features. §5.7's window and this ladder ask the same question of the same event and read the same column; a target still broken keeps producing fresh failures, and each one both re-arms this backoff and holds the target in the attention queue.

Do not search before the release exists. TMDB carries release dates; a movie with no digital release date gets zero targeted searches. This is the single largest source of wasted queries in Radarr and it is free to avoid.

6.3 Blacklist

Keyed on infohash and normalised release name. Anything that hard-failed post-ffprobe is never grabbed again, including by RSS.

A manual blocked flag stops targeted search for a title entirely while leaving RSS matching on.

7. Download and import

7.1 qBittorrent

WebUI API v2 at qbittorrent.n62.casa. Labels are qBittorrent tagsmovies-main, tv-kids and so on — enough to find things in qBittorrent's own UI, and distinct from Radarr's existing radarr/sonarr labels so both stacks can run side by side. Tags rather than a category because a category also governs the save path, and arr owns that.

torrents/add answers Ok. and nothing else: no hash, no name, and no signal that the torrent was already there. arr therefore derives the v1 infohash from the magnet or the .torrent bytes before the call, and looks the torrent up by it. That is also what makes an add idempotent across a restart (§8) — the same release resolves to the same hash, and a second add is recognised as the duplicate it is.

Verified: Radarr's only media bind is /mnt/media-v2:/mnt/media, one mount, one dataset, and the existing arrs hardlink-import today. So link() succeeds.

Implementation still falls back to copy on EXDEV rather than trusting that forever — no configuration flag, no way to get it wrong.

Hardlinking gives the library file and the seeding file independent names at zero extra disk. Copy would double storage for the entire seeding window, which at 4K is 40-80 GB per title.

7.3 Two lifecycles, deliberately decoupled

The torrent and the library entry are separate state machines.

Seeding obligation is per tracker, configured locally because Prowlarr does not expose tracker rules — ratio and min_seed_time. Set ratioLimit and inactiveSeedingTimeLimit on the torrent at add time, with shareLimitAction set to stop rather than delete, and let qBittorrent enforce them. A reaper deletes torrents qBittorrent stopped on a limit arr set.

Stopped-on-a-limit, not merely stopped: a torrent the operator paused by hand looks identical in the state field alone, and the reaper deletes data. So the ratio and idle limits are checked against the torrent's own counters, and a torrent stopped by a global limit reads as still seeding — the safe direction to be wrong in.

Consequently a hard-failed release is blacklisted and never imported, but its torrent keeps seeding until the obligation clears. Nothing is deleted early to satisfy the library.

7.4 Layout

/mnt/media/movies/main/Dune Part Two (2024) [tmdbid-693134]/
    Dune Part Two (2024) [tmdbid-693134] - [2160p][WEB-DL][HDR10].mkv

/mnt/media/tv/kids/Bluey (2018) [tmdbid-82728]/
    Season 01/
        Bluey (2018) - S01E02 - Hospital [1080p][WEB-DL][pt-PT].mkv

Media kind first, hard audience boundary second, people nowhere.

  • Provider ID in the folder name turns Jellyfin matching from fuzzy string guessing into exact lookup.
  • Attribute tags come from ffprobe, not from the release name, so they are true. The filename doubles as an audit surface: ls shows which files are DV or which of the kids' files are still English-only, with the app stopped.
  • One folder per title even for a single file, so sidecar subtitles and artwork stay contained and deletes are atomic.
  • Release group is deliberately absent. It is not a selection criterion and it makes filenames long enough to break a terminal.

Changing a title's root relocates its title folder into the new root; roots are assumed to share one filesystem, so the move is a rename, never a copy. Changing a root's path is the same move over every title under it, and it is all or nothing: one folder that cannot move puts back the ones that already did and leaves the root's path alone, so the stored path always describes the disk.

During transition, write into the existing roots so Jellyfin needs no reconfiguration and new content appears immediately. Radarr will not touch a folder it has no record of.

7.5 Jellyfin

On successful import, call Jellyfin's refresh endpoint for the affected library. Its filesystem watcher misses things. One HTTP call at the end of import.

8. The reconcile loop

There is no job queue.

The database rows are the work list. A wanted episode with no file is a pending grab. A torrent past its seeding rule is a pending delete. A downloaded file not yet probed is a pending import. Every tick, compare desired state to actual state and act on the gap.

This is idempotent and crash-safe by construction. Kill the process mid-grab and the next tick recomputes the same gap and continues. A job table would need retry counts, dead-lettering and reconciliation against qBittorrent anyway, because qBittorrent is an external system that changes underneath the app.

Transient state — a search in flight, download progress — is in memory and rebuilt from qBittorrent on startup. Where that state is ever seen is §9.8: inline on the row that owns the item, never persisted.

Ticks are staggered: reconcile every 30 s, RSS every 10 min, metadata refresh daily, reaper every 5 min.

9. API and UI

9.1 API-first

The HTTP API is the product surface; the web UI is a client of it with no privileged path. OpenAPI spec generated from handler annotations, served alongside the app, and used to generate the TypeScript client.

Own schema first. Radarr-compatible endpoints are a later, separate, lower-priority crate (§9.4), not the primary shape.

One box. Two grouped result sets: in library first, on TMDB below. Both sets are titles — series and movies. Episode rows never appear, under any query, and there are no season rows. Enter on a TMDB result opens the add flow with root and policy pre-filled.

The same box accepts a raw TMDB or IMDb ID, and a pasted magnet or .torrent, which skips to the manual-grab flow.

There is never a moment where the user has to know whether they are searching or adding.

Result rows are enriched: a poster thumbnail and a rating, under the same rules as title detail (§9.6) — images hotlinked from path fragments, the rating being TMDB's vote_average with its vote_count. A trailer chip renders per row and resolves only when clicked (§9.6).

9.3 Manual search results

Radarr's manual search is unusable because the raw release name is the dominant column, pushing everything that actually decides the choice off-screen.

Invert it. The policy engine has already classified every candidate:

  • Three buckets. eligible shown by default, sorted by score. waived (fails a soft rule, grabbable with one click that writes an override) and rejected collapse to a count, expandable.
  • Columns are parsed attributes as chips — score, resolution, source, HDR, audio languages, size, seeders. Fixed width, no horizontal scroll.
  • Release name is secondary, truncated, full string on expand. It is evidence for when you disagree with the parse, not the primary key.
  • Every rejected row names the rule that killed it, so over-strict filters are visible without reading names.

Pre-grab, HDR and audio chips are best-effort from the name (§5.6).

9.4 Jellyseerr compatibility

Jellyseerr stays. It already does Jellyfin user auth, discovery and request approval — none of which is the problem being solved, and rebuilding it doubles the project.

A thin arr-compat crate exposes the slice Jellyseerr actually calls:

GET  /api/v3/system/status
GET  /api/v3/rootfolder            → the real roots
GET  /api/v3/qualityprofile        → one fake profile per root
GET  /api/v3/tag
GET  /api/v3/movie                 existence check
POST /api/v3/movie                 add
GET  /api/v3/movie/lookup

plus the series equivalents. Thin mappings onto the real domain, isolated in one crate that never leaks into arr-core.

Lower priority than everything in §5-8.

9.5 Notifications

ntfy, one topic per Owner. Three events only — Radarr's failure mode is notifying on everything and being muted within a week.

  • Imported → to the title's owners. The only good-news notification.
  • Needs a decision → to the operator alone. Entered the no-PT-source queue, or the needs-a-decision queue (§5.7).
  • Broken → to the operator alone. Prowlarr, qBittorrent or TMDB unreachable, disk full.

Not notified: grabs, searches, downloads starting or finishing, soft fails.

9.6 Title detail

One detail surface per kind: /movies/{id} and /series/{id} (#129). For a movie the release deck becomes a section of the page, and /movies/{id}/releases keeps resolving; series keeps its season-and-episode shape, seasons ordered newest-first within the series and episodes newest-first within each season — the seasons you are deciding about now are at the top, not after nine rows of back catalogue.

TMDB is the only metadata source. The rating shown anywhere is TMDB's vote_average with its vote_count, rendered as amber stars (--signal-warn — the same state colour a wanted chip carries). No OMDb, no IMDb or Rotten Tomatoes scores — each would need a second upstream, a second key and a second thing that can be down.

Images are hotlinked from image.tmdb.org. The API returns TMDB path fragments, never URLs; the browser composes the URL and chooses the size. No image proxy and no image cache in the service.

Rich detail is not persisted. It is served through arr-meta's 24-hour response cache, which lives on disk in a directory alongside the database (§10), bounded by both age and total size (#158) — not the database itself, and not an image cache; images stay hotlinked as above. The single exception is poster_path, backdrop_path and vote_average, stored on movies and series and written by the daily metadata refresh (§8), so library views render without a TMDB call.

External links are TMDB always, IMDb for movies, TVDB for series — all from ids the app already holds — plus a Rotten Tomatoes search link, which is a query URL, not a resolved title page.

Trailers resolve on click. TMDB's search responses carry no videos, so a trailer key costs a detail call. Rendering one chip per title and resolving the one clicked keeps that cost at one call, and the 24h cache makes a repeat free.

Library view. A poster grid by default with a list toggle; the list keeps the derived-status columns §4.2 built it around.

Out of scope here, because they are the adjacent scope most likely to creep: watch providers, recommendations or similar titles, collections, person pages inside the app, review text.

9.7 The shell

The library is the homepage. / renders the library view described in §9.6. It is what the operator opens the app to look at, so it is what the app opens on.

The signal chain is a settings section. The four upstreams — tmdb, prowlarr, arr, qbittorrent — and their lamps live inside /settings, and have no route of their own. Per-upstream health is something you check when something is wrong, not a homepage.

The master lamp stays on the rail, so the at-a-glance read that the chain is healthy survives the move and no navigation is needed to get it.

The wordmark is a link home. arr on the rail navigates to /.

The rail composition is fixed. From left to right it carries the master lamp and wordmark, a warning mark only while queues need attention, the centred search, the version readout and the settings cog.

9.8 Downloads on the row

§8 keeps download progress as transient state; nothing in §9 so far says where it is seen. Without a rule here, a season pack downloading and one that never started look identical.

Progress is inline, on the row that owns the item — a movie row, a season row for its pack, an episode row. There is no downloads page and no count in the rail: a download is an attribute of the thing being downloaded, not a place to visit.

Alongside progress the same row carries the other states a torrent can be in: seeding under §7.3's obligation, stalled, errored.

A torrent arr did not grab is never shown. §2 already rules the service knows only what it put on disk; qBittorrent's own UI lists the rest.

Nothing is persisted — no progress column, no new table. The snapshot comes from qBittorrent and dies with the process (§8), refreshed by the UI's normal polling cadence at roughly 15 s rather than SSE.

At phone width only active-grab progress survives; seeding and stalled shed with the row's other secondary chips.

Out of scope here: any change to what the reconcile loop does, and any new state machine — §7.3's two lifecycles stay as they are.

10. Persistence

SQLite via sqlx, compile-time-checked queries, migrations in arr-db.

A few thousand rows, single writer. Postgres would buy nothing and cost a service.

Policy lives in the database, not a config file — size targets and DV rules get tuned by hand during testing and a restart-to-reload loop gets old immediately. Only bootstrap settings (bind address, Prowlarr URL, qBittorrent URL and login, TMDB key, media root, TMDB response cache directory) come from config/env. The qBittorrent password is a secret, so it is env-only and has no config-file field.

Backup is sqlite3 .backup on a timer.

11. Crate layout

Cargo workspace, members = ["crates/*"], versions pinned once in [workspace.dependencies], following ~/tea/maestro.

arr-core      domain types, policy engine, scoring     no IO, no heavy deps
arr-parse     release name parsing                     no IO
arr-meta      TMDB client
arr-indexer   Torznab via Prowlarr
arr-dl        qBittorrent WebUI API
arr-probe     ffprobe wrapper
arr-subs      subtitle providers, translation, sync
arr-db        sqlx + migrations
arr-api       axum + OpenAPI
arr-compat    Radarr/Sonarr v3 shim for Jellyseerr
arr-daemon    reconcile loop, wires everything
arr-e2e       cross-process integration tests
web/          Vite + TypeScript SPA, embedded via include_dir

The rule: arr-core and arr-parse hold the logic worth testing constantly and must not depend on axum, sqlx or reqwest. Everything expensive is downstream of them.

Nothing is generic over media kind. Movies are built concretely, then TV concretely, and shared machinery is extracted only once both exist. An abstraction derived from one example fits one example.

12. CI

Lints declared once in the root manifest:

[workspace.lints.rust]
unused_crate_dependencies = "warn"
missing_debug_implementations = "warn"

[workspace.lints.clippy]
pedantic = { level = "warn", priority = -1 }
unwrap_used = "warn"

Member crates carry only [lints] workspace = true.

Per-push gate, target under 5 minutes:

Step Tool
format cargo fmt --check
lint cargo clippy --all-targets -- -D warnings
unused deps cargo machete (stable; cargo udeps needs nightly and a full rebuild)
test cargo nextest run
frontend biome ci web/ and tsc -b --noEmit

Off the gate, scheduled: cargo deny for advisories and licenses. Coverage, if ever, likewise. Neither blocks a push.

Total wall-clock on the self-hosted Gitea runner is an explicit constraint. The lever is caching, not step selection: cache ~/.cargo/registry, ~/.cargo/git and target/, keyed on Cargo.lock plus rust-toolchain.toml. Never build --release in the gate.

End-to-end tests, only at the seams that actually break:

  • Prowlarr and TMDB — wiremock with recorded real responses as fixtures. Never live: trackers rate-limit, and it would leak credentials into CI.
  • qBittorrent — a real container, started with a seeded config because the image prints a random WebUI password per boot. Its API semantics are the most likely source of surprise: torrents/add reports nothing back, and required parameters have changed between major versions.
  • ffprobe — tiny committed clips, a few KB each. The Jellyfin LXC already has the right fixtures at /srv/jellyfin-test, including a real DV Profile 5 clip. That is the test proving the policy engine rejects Profile 5 and accepts 8.1 — the subtlest rule in the system.

E2E runs on main and on pull requests touching those crates, not every push.

13. Build order

Each phase ends at something usable end to end. No phase is a refactor of the previous one.

  1. Skeleton — workspace, CI gate, config, SQLite migrations, health endpoint, embedded empty SPA.
  2. Parse and scorearr-parse and arr-core against fixture release names. Pure, fast, heavily tested. No network.
  3. Read-only sourcing — TMDB lookup, Prowlarr enumeration and search, classified results over the API. Still grabs nothing.
  4. Movies, end to end — add a movie, grab, download, ffprobe, hardlink, rename, Jellyfin refresh, seeding reaper. The first real cutover test is one movie Radarr does not know about.
  5. UI — unified search, manual search buckets, library views, the queues.
  6. TV — seasons, episodes, auto_track, per-episode versus season-pack grabbing, derived status.
  7. Owners and notifications — tags, per-person ntfy topics, filtered views.
  8. Jellyseerr compatarr-compat.
  9. Subtitles — replaces Bazarr. See §15.

Movies before TV because TV adds season packs, air-date calendars and per-episode state on top of an otherwise identical pipeline. Doing it second means that pipeline is already proven.

14. Open questions

  • Remux playback. Whether a 4K remux streams cleanly to the Shield is untested. The usual failure is audio, not bitrate: BluRay remuxes carry TrueHD/DTS-HD MA, and if the Shield cannot bitstream that downstream, Jellyfin transcodes audio and playback stutters while video direct-plays. Test before fixing the 4K size ceiling in §5.5. The outcome changes the conclusion from "remuxes are bad" to "remux audio needs a downmix".
  • Size band numbers in §5.5 are placeholders pending that test.
  • Season-pack re-grab. When an airing season completes, the episodes are already present individually. Nothing re-grabs the pack, and per §6.2 the RSS lane obeys the same guard — it skips packs for any season with episodes on disk. Whether a re-grab is ever wanted remains unresolved and deliberately deferred.

15. Subtitles

Replaces Bazarr. Phase 9 in §13; the arr-subs crate in §11.

Wanted set. Global, not per root. Two languages are separately wanted for every media file: Portuguese — pt-PT preferred, pt-BR accepted — and English. A file is satisfied for a language when a subtitle in it exists, embedded or as a sidecar. Satisfying a language is a statement about viewing it, and not about having text in it — the distinction matters where an image track is all there is, and Translation below is where it bites. This is deliberately unlike §5.2's audio rules, which attach to a root: subtitles carry no blacklist there and none here. pt-BR subtitles are always fine.

Embedded tracks. An embedded subtitle track satisfies its language. Text-format tracks (subrip, ass, mov_text) are additionally extracted to a sidecar SRT, because an extracted track is a legal translation source. Image-format tracks (PGS on BluRay, VobSub on DVD) carry bitmaps, not text: they satisfy viewing but can never feed a translator, and arr does not OCR them. arr-probe already reports subtitle tracks with resolved languages; the format is the new fact it must carry.

Cue skeletons. What an image track does have is exact timings, because its packet timestamps are the disc's own cue structure. At import — while the file is being read and hardlinked anyway — arr derives a cue skeleton from each non-forced image track: ffprobe -show_packets on the track, packets paired show-to-clear, no pixel read. The skeleton has timings and no text, and that is enough, because alass matches on interval structure rather than on words. A sparse skeleton is still a strong reference; a downloaded subtitle and a retail disc's track never carry the same cues anyway.

The pairing is the whole of it, so it is validated before it is trusted: an even packet count, every implied duration plausible, and a cue density that fits the runtime. PGS permits several composition segments per subtitle, and a track built that way pairs into something that is quietly half a second out — worse than no skeleton. A track that fails validation gets none, and alignment falls back to the video.

Providers. OpenSubtitles.com, behind one trait.

Ranking. A moviehash match wins outright. Then an exact release-name match, then same release group or same source, then uploader rating and download count as tiebreakers. Every rejected candidate names the rule that killed it, so the manual view described in §9.3 works unchanged for subtitles.

Forced and SDH. A forced track covers only foreign-language lines and on-screen signs; it never satisfies a want and arr never goes looking for one. An SDH track is complete and satisfies, ranked below a plain subtitle.

Neither gets its own sidecar name. A language is satisfied by exactly one sidecar, so <video>.<lang>.srt needs no segment distinguishing forced from plain from SDH — where a plain subtitle exists for a language, the forced one is ignored rather than kept beside it. The database says the same thing: one sidecar row per (media file, language).

Translation. When no provider has a wanted language, arr translates immediately — there is no waiting window. The source is an existing subtitle: a downloaded one, or one extracted from a text-format embedded track. Being able to translate from an embedded track is a deliberate improvement on Bazarr, which cannot.

Fetching a source. A release whose only subtitle is an image track has no text source and never will: the track satisfies its language, so nothing is ever fetched in it, and the track itself cannot be translated from. That is a deadlock, and it is broken by a narrow carve-out — when a wanted language needs translating and no text source exists, arr fetches a text subtitle in the image track's language even though that language reads as satisfied. The fetch obtains a source; it settles no want of its own, and the no-upgrade rule below does not apply to it. Where that track has a cue skeleton, the fetched subtitle is aligned against the skeleton rather than the video, which puts it on the disc's own timings instead of on a heuristic read of the audio. OCR remains a non-goal: this reaches the same place with real text.

Translation backends. Pluggable, each behind its own cargo feature: an OpenAI-compatible HTTP endpoint, DeepL, Google Translate, and a generic remote command driven by a configured template (ssh box claude -p is one instance of that template, not a backend of its own). The engine in use is a database setting, so switching does not need a rebuild when the feature is compiled in. Subtitles are sent in batches of cues; a reply whose cue count or numbering does not match the batch is rejected. Timing data never leaves arr.

No upgrade loop. Once a language is satisfied — by a machine translation too — arr stops working on it. A real subtitle appearing later does not replace anything. Replacement is a manual action from the UI. This is §5.4's rule applied to subtitles.

The one exception is the source fetch above. It is not an upgrade: the language it downloads is already satisfied and stays satisfied by the same track it was before, and what the download settles is a different language's gap. Nothing is replaced, so nothing about the rule changes.

On disk. Sidecars live next to the video inside the §7.4 title folder, named <video basename>.<lang>.srt, e.g. … - [2160p][WEB-DL][HDR10].pt-PT.srt. A machine translation carries an extra .mt segment: … [HDR10].pt-PT.mt.srt. That keeps §7.4's audit-by-ls property — with the app stopped, the filename says which subtitles are machine-made. Folder-level delete stays atomic because sidecars are inside the folder.

.mt is the only optional segment. One language, one sidecar: a second subtitle for a language arr already has is refused, and replacing one is the manual delete-then-fetch §15's no-upgrade rule already describes.

Sync. alass runs on every fetched and every translated subtitle. It is a single small binary invoked like ffprobe, so it costs nothing at rest. It reports no confidence value, so its output is accepted unless it is implausible — a shift beyond 60 seconds, or cues lost — in which case the unsynced original is kept and the file is flagged.

The reference is the video, except where a cue skeleton exists, and then it is the skeleton. A subtitle a skeleton accepted is already on disc-exact timings, and translation copies those timings over untouched, so the pass that would otherwise run on the translation is skipped: a second alignment has nothing left to find and can only move them off.

Configuration. Provider credentials and translator API keys are bootstrap config or environment, per §10 — a secret never becomes a database row. Wanted languages, chosen engine, per-provider enable and the daily budgets are database rows edited from /settings without a restart.

So is everything needed to point the OpenAI-compatible backend somewhere else: its base URL and its model name are database rows too, not bootstrap config. That backend is not "OpenAI" — it is any endpoint speaking that shape, llama.cpp and a local gateway included, and which one is in use is a thing to try and change, not a property of the deployment fixed at start-up. Only the API key stays in the environment, and an endpoint that needs no key is a valid configuration.

Budgets. A token bucket per provider and per translator, with a configured daily allowance. The reconcile loop spends it newest-import-first, so enabling this on an existing library drains the backlog over days instead of hitting every rate limit at once. Being at the cap is a visible queue state, not an error.

Notifications. No new event classes. §9.5 stands: subtitle fetches never notify, and a provider or translator being unreachable folds into the existing "Broken" message to the operator alone.

Non-goals, stated here so they do not creep back: OCR of image-based tracks, transcribing audio when no subtitle exists anywhere, adopting subtitle files already on disk that arr did not write (§2 already says the service knows only what it put there — which does mean arr may fetch a second copy alongside one Bazarr left), and any background loop that upgrades a subtitle in place.