A release whose only subtitle is a PGS or VobSub track was stuck both ways: the track satisfied English so no English SRT was ever fetched, and bitmaps can never feed a translator. §15 is amended to separate satisfying viewing from providing a translation source, and to carve a fetch made to obtain a source out of the no-upgrade rule. Timings come from the disc rather than from alass guessing at the audio. At import, each non-forced image track's packet timestamps are paired show-to-clear into a cue skeleton and stored; the fetched source is then aligned against that skeleton, and the translation made from it skips the post-translation pass, which could only move disc-exact timings off. Pairing is validated before it is trusted — even packet count, plausible durations, sane density for the runtime — because PGS allows several composition segments per subtitle and a slipped pairing is quietly half a second out. A track that fails validation gets no skeleton and falls back to aligning against the video, as does any file imported before this: there is no backfill. Closes #268 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
50 KiB
arr — design
One service that replaces Radarr and Sonarr, and later Bazarr, for a single household. Rust workspace, API-first, SQLite, no authentication layer.
This document is the contract. Issues reference it by section rather than restating it.
1. Why
Five .NET services idle at 600-900 MB to do a job with one loop in it: decide what you want, find it, fetch it, put it somewhere Jellyfin can read. Each has a configuration surface an order of magnitude larger than the subset actually used, and none of them share a model, so the same show is configured three times in three dialects.
The replacement targets 20-40 MB resident, one binary, one database, one policy language.
This is not a general-purpose Radarr replacement. Every decision below is allowed to be narrow.
2. Non-goals
Explicitly out of scope, permanently unless stated:
- Authentication. The perimeter is a VPN plus Authelia at the proxy. The service binds without auth, same trust model as the rest of the stack.
- Library migration or filesystem scan. The service knows only what it put on disk. Adopting the pre-existing library, if ever wanted, is a one-off script against both APIs, not a feature.
- Automatic quality upgrades. See §5.4.
- Multiple quality profiles. Policy attaches to a root folder. See §5.1.
- Trakt. Jellyfin and Stremio already write watch state there and nothing in this pipeline reads it.
- Calendar view, actor/director following, per-indexer UIs. Prowlarr owns indexer configuration and keeps doing so.
- Books and music.
shelfmarkandcalibre-web-automatedalready cover books. If they ever land here they reuse maybe three of the ten crates, so nothing is generalised in advance. See §11.
3. System context
┌──────────────┐
TMDB ──────▶│ │
│ arr │────▶ qBittorrent WebUI qbittorrent.n62.casa
Prowlarr ──────▶│ │
(Torznab) │ (this) │────▶ ffprobe local subprocess
│ │
Jellyseerr ─────▶│ │────▶ Jellyfin refresh 10.6.10.18:8096
(Radarr-compat) └──────┬───────┘
│ ntfy per-person topics
▼
SQLite + media tree on /mnt/media
Everything already exists except arr. Prowlarr keeps owning tracker auth,
Cloudflare bypass via FlareSolverr, rate limiting and the Cardigann
definitions — replacing it buys nothing.
qBittorrent is reached at qbittorrent.n62.casa, WebUI API v2, download dir
/mnt/media/qbittorrent/complete. Its WebUI requires a login, so arr carries
credentials — the one upstream that does.
4. Domain model
Movies and series are separate aggregates. They share machinery, not a table.
Root
id, kind (movie|tv), audience (main|kids), path, policy_id
Policy
id, name
required_audio rule set, see §5.2
dub_blacklist [pt-BR]
hdr_rules see §5.3
size_bands per resolution: floor, target, penalty curve,
all per episode (§5.5)
resolution_pref [2160p, 1080p]
source_weights small tiebreaker
Movie
id, tmdb_id, title, year, original_language, root_id
wanted (bool), overrides (json), state, blocked (bool)
search_attempts, last_searched_at
Series
id, tmdb_id, title, year, original_language, root_id
auto_track (bool), overrides (json), blocked (bool)
Season
id, series_id, number, tracked (bool)
Episode
id, season_id, number, title, air_date
wanted (bool), state
search_attempts, last_searched_at
MediaFile
id, owner_kind (movie|episode), owner_id, path, size
probed (json: resolution, source, hdr, audio_tracks, sub_tracks)
waiver (json|null) which rule was relaxed to allow this import
Release
id, indexer_id, guid, name, size, seeders, publish_date, download_url
parsed (json) what the name claims
score, verdict eligible | waived | rejected(rule)
Grab
id, release_id, target_kind, target_id, infohash
state, grabbed_at, imported_at
Blacklist
infohash, normalised_name, reason, created_at
Owner
id, name, ntfy_topic
TitleOwner
title_kind, title_id, owner_id many-to-many
4.1 Intent lives at the leaf
Sonarr's monitored conflates "do I want this show" with "is anything
happening right now". It is user-maintained, so it drifts and stops meaning
anything.
Here, intent is Episode.wanted and Movie.wanted only.
Series.auto_track is not intent — it is a rule one level up: when metadata
reveals a new season, that season becomes tracked. The rule applies from the
second refresh onward: seasons revealed by a series' first refresh — the whole
back catalogue present at add time — are never tracked by it. Only seasons that
appear after the series was added are covered; back-catalogue seasons are picked
by hand. An untracked series where you
manually marked S02 needs no special case: three wanted episodes, nothing else.
Season.tracked is also a rule, not intent. Turning tracking on marks every
already-revealed episode in the season wanted, and keeps marking episodes
wanted as metadata reveals them. Turning it off clears wanted on every
episode of that season: toggling the flag is itself the explicit act, and the
product offers no other way to withdraw a season's worth of intent. Without
that, untracking leaves a series with every season off and every episode still
wanted, pinned at incomplete with no way back. Marking individual episodes
wanted on a season that was never tracked is unaffected by any of this.
Both rules skip season 0. Specials are dozens of undated shorts and recaps
that indexers do not carry, so auto_track never marks season 0 tracked. See
§4.2 for how season 0 stays out of derived status.
4.2 Status is derived, never stored as intent
Recomputed on metadata refresh and on file change:
| Status | Condition |
|---|---|
airing |
an episode aired or airs within ±14 days |
incomplete |
wanted episodes missing, nothing currently airing |
waiting |
season finished, next unannounced or future-dated |
complete |
everything wanted is on disk |
ended |
series finished upstream and complete |
Default list view shows airing and incomplete. Everything else collapses
behind one toggle. A one-off season grab therefore disappears from the default
view by itself once satisfied, and a tracked show reappears by itself when a new
season is announced. Nothing to remember to flip.
Status chips carry state colour, not a single informational hue: green
(--signal-ok) for on disk or complete, amber (--signal-warn) for wanted but
not yet found, violet (--status-airing) for downloading, neutral for
untracked. The word stays present in every chip, so state survives with colour
removed.
One state is exempt: a file being on disk is drawn as a check glyph rather
than the word available, with the word on the chip's aria-label. A drawn
check is a shape, not only a hue, so it survives the greyscale read the rule
exists to protect. The exemption is that narrow — every other state, and every
derived-status chip above, keeps its word, because no unambiguous glyph stands
in for incomplete or waiting.
Season 0 is invisible to all of it: derived status ignores season 0 episodes
entirely, so a manually wanted special cannot pin a series at incomplete or
hold back ended. The accepted consequence is that specials are visible and
grabbable in the series detail view, but never affect status.
4.3 People are tags, not paths
Owner is a many-to-many tag on titles. It drives UI filtering, the default
list, and notification routing. It never appears in a path. A file lives in one
place; audiences overlap.
5. Policy engine
The core of the app and the part worth testing hardest. Lives in arr-core,
pure, no IO.
5.1 Policy attaches to a root
Two roots per media kind (main, kids), each with one policy. Jellyseerr's
"pick a quality profile per request" model is satisfied by exposing exactly one
fake profile per root and ignoring whatever it sends.
Per-title overrides relax the root policy for one title. Same mechanism in
both directions — only_4k tightens, allow_english_audio loosens.
Stored verdicts follow the effective policy. A release's verdict is stamped by the search that found it, and both §9.3's deck and the daemon's manual-grab gate read that stored column. So the four operator actions that change a title's effective policy — editing its overrides, moving it to a root with a different policy, pointing a root at a different policy, and editing the contents of a policy some root points at — re-derive the stored verdicts of everything they touch, inline in the same request. The operator is never left reading a verdict computed under a policy that no longer applies.
The fourth is the widest: a repoint moves one library, a policy edit moves every library sharing the policy. Inline still holds there. Re-evaluation is pure and in-memory, a row whose verdict does not move is not rewritten, and the ceiling is the whole database rather than something that grows with the number of roots — roots partition titles, and a title has exactly one root. Measured over 2000 titles and 10 000 stored releases split across two roots sharing one policy: 0.36 s when the edit moves no verdict, 2.7 s in the pathological case where it moves all 10 000. A rename changes no rule and re-derives nothing.
5.2 Language
Requires a concept the release name does not carry: the original language of the title, from TMDB, distinct from a track's language.
reject audio track T if T.lang ∈ dub_blacklist and T.lang ≠ title.original_language
dub_blacklist = [pt-BR] globally. That single expression covers both roots:
- main — require a track matching
original_language. Others may ride along. A Brazilian film hasoriginal_language = pt-BR, so its own soundtrack passes; a pt-BR dub of an English film does not. - kids — require pt-PT, or
original_languagewhen that is already Portuguese. pt-BR never satisfies the requirement, because the child does not read and a Brazilian dub is not what he should be watching. - subtitles — no blacklist. pt-BR subtitles are always fine.
The pt-PT / pt-BR detection problem. ISO-639-2 has one code, por, for
both. ffprobe reports por either way. Three signals in order:
- Release name markers — Brazilian releases self-identify loudly:
PT-BR,Dublado,Nacional,Dual Áudio. ffprobestream title andhandler_namestrings, not the language code — often literallyPortuguese (Brazil).- Container-level BCP-47 tags where present (Matroska can carry
pt-BR).
When none resolve it, tag the track por-unverified and surface it. Do not
guess. Guessing toward kids gives a child a Brazilian dub; guessing toward
main rejects a Brazilian film's own audio.
Sourcing consequence for kids. European Portuguese dubs live almost
entirely in streaming WEB-DLs (Disney+, Netflix, Max) and essentially never in
BluRay encodes or remuxes. That root biases hard toward WEB-DL sources, and
unlike audio in general it is often pre-grab visible, since multi-audio
releases advertise themselves (MULTi, DUAL, explicit language lists).
When nothing qualifies, the title sits in a visible no-PT-source queue
rather than being grabbed in English. That queue is acted on manually,
sometimes for months. A per-title allow_english_audio override empties one
item from it with one click, and the resulting file is recorded with a
waiver so the UI shows it as "English, no dub" rather than a clean match.
5.3 HDR and Dolby Vision
The naive rule "blacklist Dolby Vision" is wrong and would discard a large share of good 4K releases.
| Profile | Base layer | Behaviour without DV support | Verdict |
|---|---|---|---|
| 5 | none (ICtCp) | green/purple, unwatchable | reject |
| 7 | dual-layer BL+EL | depends on player, unreliable | reject |
| 8.1 | HDR10 | plays as HDR10 | accept |
The existing library is already almost entirely Profile 8.1 with no Profile 5, and Jellyfin tonemaps it correctly on the current GPU path.
Profile is knowable only from ffprobe, never from a release name. A DV chip
in search results means "the name claims DV" and its absence means nothing.
5.4 Quality, and why there is no upgrade loop
Default: take 4K if it exists now, otherwise 1080p. Per-title only_4k
rejects 1080p outright and keeps searching indefinitely — the escape hatch for
the handful of titles worth waiting on.
No grab delay. Deliberately rejected. Radarr's delay profile exists to stop
first-match-wins from always yielding 1080p; here the only_4k flag solves the
same problem upfront and without making every grab slower.
No automatic upgrade loop. An automatic upgrade re-downloads 40-80 GB and
restarts a seeding obligation for something already watched. needs_upgrade is
a marker in the UI and a manual "search again" button, not a background job.
5.5 Scoring: distance from a target size
Source tier is not the dominant term. A remux should win only when nothing smaller exists.
Per resolution, a floor, a target and a growing penalty above target. For 4K: floor ~8 GB, target ~22 GB. A 20 GB WEB-DL sits at target and wins; a 60 GB remux scores badly but stays eligible, so it is picked when it is the only option. The floor matters — unbounded "smaller is better" selects a 3 GB 4K encode that looks like mud.
A band describes one episode. For a movie the question never arises — one release is one film. For TV a release may carry a single episode, several, or a whole season, and the shipped TV numbers are episode numbers: a 1080p target of 2 GiB is what one episode should weigh, not a ten-episode pack. So a release is measured by its size divided by the number of episodes it covers, and that per-episode figure is what both the target penalty and the floor are compared against. A 20 GiB pack of ten episodes is scored as 2 GiB and sits at target, not ten times above it.
The floor takes the same figure. It is a hard reject rather than a penalty, so comparing a pack's total against an episode-sized floor lets every pack through untested — wrong in the opposite direction from the target penalty. The floor exists to keep a 3 GB 4K encode out, and a pack of ten such encodes has to fail it just as plainly.
An unknown episode count is one episode. A release matched before its season's episodes have been revealed by a metadata refresh has nothing to divide by, and the rule does not guess a number: the divisor is 1 and the release is measured by its full size. That is the behaviour today, and it makes a pack look oversized rather than undersized — it fails toward rejecting a good pack rather than grabbing a bad one, and the next search after a refresh has the real count.
A band describes a rate, not a fixed size per episode. The shipped values
are read against a reference runtime of 45 minutes: a 1080p target of 2 GiB
means 2 GiB per 45 minutes of episode. Before the per-episode figure is compared
against them, a band's floor and target are both scaled by
runtime / 45, where runtime is the series' minutes-per-episode from TMDB
metadata. For a typical drama around 45 minutes the shipped numbers keep exactly
their current meaning; a 22-minute show is judged against roughly half the floor
and half the target instead of being rejected for weighing half of what an hour
of video weighs. The shipped band values themselves do not change — the
reference runtime is chosen so they do not have to.
A missing or zero runtime is the reference runtime. When metadata carries no per-episode runtime, or carries zero, the scale factor is 1 and the band applies unscaled — exactly today's behaviour. The rule does not guess a duration, for the same reason the unknown episode count does not guess a number.
Movies are not scaled. This applies to episodes only. A movie's bands are already tuned against feature length, so its floor and target keep their current meaning regardless of the movie's own runtime.
Source tier (Remux > BluRay > WEB-DL > WEBRip > HDTV) survives as a small
tiebreaker. Seeders are log-scaled and small: enough to complete, past that it
does not matter. Telesync, CAM and screener are hard filters, not low
scores.
How resolutions rank against each other. Size is scored against the band
for the release's own resolution, so on its own it says nothing across
resolutions: an at-target 4K and an at-target 1080p both score the top of the
size term, and a 4K a couple of gigabytes over target loses to a 1080p that is
merely on target. That is wrong — resolution_pref is an ordered list, and the
order is a preference, not just an eligibility filter.
So each step up resolution_pref is worth a fixed number of points. The last
entry is worth nothing and every earlier one a step more. A resolution the list
does not carry scores nothing rather than being penalised, the same as an
unclaimed resolution: no ranking is no opinion.
The step is sized against the size term, not chosen in isolation. With the seeded 4K band — target 22 GB, 60 points per gigabyte over — a step of 300 is five gigabytes of overshoot: a 4K up to about 27 GB beats an at-target 1080p, and a bloated 40 GB 4K does not. Below target the same arithmetic asks a 4K to be within about four gigabytes of its target to win, which is what keeps an 8 GB 4K that looks like mud from beating a good 1080p.
Exact numbers are policy rows, tuned by hand. The model is the decision.
5.6 Two phases of truth
Release names lie or omit. Files do not.
- Pre-grab — parse the name. It is all there is. Use it for hard filters (source type, resolution, explicit language markers) and for scoring.
- Post-download —
ffprobebefore import. Real audio languages, real HDR format and DV profile, real codec, real duration.
Store both, parsed on the release and probed on the file. The gap between
them is also data: it tells you which release groups mislabel.
5.7 Hard fail versus soft fail
A policy violation found by ffprobe is not one thing.
- Hard fail — useless. No original-language audio track at all, DV Profile 5, wrong title, corrupt. Do not import, blacklist the release, grab the next candidate.
- Soft fail — watchable but not what was asked. 1080p while
only_4kwas set, PT audio missing on akidstitle. Import it, record awaiver, surface it.
Neither deletes the torrent. See §7.3.
Two hard failures make a decision, and only for 30 days. A movie, an episode or a season enters the needs-a-decision queue (§9.5) when two grabs against different releases hard-failed on it, and both of those failures happened within the last 30 days. One bad torrent is not a decision — a release that hard-failed is blacklisted (§6.3) and the next candidate is grabbed, which is the system working.
The window runs from the failure, not the grab. The two are usually minutes
apart, but a torrent can sit stalling on a dead swarm for five weeks before
ffprobe finally condemns it — and that failure is fresh evidence the target
is broken now, not history. Measured from the grab it would be born outside
the window and a genuinely broken target could never surface. So grabs
records failed_at alongside grabbed_at, and the window reads it.
The window is what lets the queue be emptied. Nothing clears a grabs row, so
without it the queue only ever grows and the one season that wants attention
sits behind eight that were dealt with months ago. It is the queue's version of
§6.2's "it never gives up entirely, it goes quiet": a target the operator has
dealt with stops producing failures and drops out once the last one ages past
30 days, while a target that is still broken keeps producing them — the pack
guard retries at worst weekly (§6.2) — and stays queued for exactly as long as
it is genuinely broken. Nothing is dismissed by hand and no acknowledgement
state is stored, so there is no second thing to keep correct.
Only a target still waiting for a file is queued. The count and the
window already express one rule — bounded attention — and this is another
face of it: a failure history queues a target only while that target still
has a gap to fill. A movie or an episode is queued while it is wanted and
not available. A season holds no intent of its own (§4.1), so it is queued
while at least one of its episodes is still wanted and not available. A
season pack that hard-failed twice, fell back to per-episode grabbing
exactly as §6.2 says it should, and was then fully acquired leaves at once
rather than waiting out the 30 days — that is the system working, not a
decision. A target that is still broken keeps producing failures and stays.
The same rule applies on all three lanes and in both readers.
GET /api/queues/attention (§9.3) and the ntfy notification (§9.5) are two
views of one queue; filtering differently tells the operator two different
stories on two channels.
6. Sourcing
6.1 Prowlarr, per-indexer Torznab
GET /api/v1/indexer enumerates. Each indexer is then addressed at its own
Torznab endpoint:
GET /{indexerId}/api?apikey=…&t=caps
GET /{indexerId}/api?apikey=…&t=search&q=…
GET /{indexerId}/api?apikey=…&t=movie&imdbid=…
GET /{indexerId}/api?apikey=…&t=tvsearch&tvdbid=…&season=&ep=
Prowlarr's aggregate /api/v1/search is deliberately not used: Torznab is the
stable standard rather than an internal UI API, it gives real RSS semantics, and
per-indexer addressing keeps indexer identity in hand — which the per-tracker
seeding rules need anyway.
t=caps per indexer records which search modes each tracker supports. Not all
do ID-based search; some are text-only.
6.2 Three triggers, three cost profiles
| Trigger | Cost | Cadence |
|---|---|---|
| RSS sync | one call per indexer, independent of library size | ~10 min, never backs off |
| Targeted search | one call per wanted item per attempt | exponential backoff |
| Manual | one call, user-initiated | on demand |
RSS is an empty-query Torznab call whose results are matched against the wanted list locally. Because its cost does not scale with the wanted list, every wanted item is matched against every RSS result forever — with one guard: the RSS lane skips a season pack for a season that already has episodes on disk, the same guard §14 applies to re-grabs. Both lanes agree, so a season does not behave differently depending on which lane sees a release first. A pack for a season with nothing on disk stays eligible. When a single episode cannot be found on its own, the escape hatch is the season release deck, not a lane exception.
Targeted search backs off 1h → 6h → 1d → 3d, capped at 7d, reset when the
title's metadata changes. It never gives up entirely, it goes quiet.
The ladder runs from the failure, not the grab. A failed season-pack grab
quiets the pack lane on that same curve, and the rung is measured from the
moment the grab entered failed — grabs.failed_at, the column §5.7's
attention window reads — not from when it was sent. The two are usually
minutes apart, but a torrent can stall on a dead swarm for five weeks before
ffprobe condemns it at import. Measured from the grab, the whole ladder has
already elapsed by the time the failure lands, so the lane retries the source
that just failed at once, which is the one thing the backoff exists to
prevent. The ladder's job is to stay off a source that has recently failed,
and "recently" can only mean recently failed.
One anchor covers both features. §5.7's window and this ladder ask the same question of the same event and read the same column; a target still broken keeps producing fresh failures, and each one both re-arms this backoff and holds the target in the attention queue.
Do not search before the release exists. TMDB carries release dates; a movie with no digital release date gets zero targeted searches. This is the single largest source of wasted queries in Radarr and it is free to avoid.
6.3 Blacklist
Keyed on infohash and normalised release name. Anything that hard-failed
post-ffprobe is never grabbed again, including by RSS.
A manual blocked flag stops targeted search for a title entirely while
leaving RSS matching on.
7. Download and import
7.1 qBittorrent
WebUI API v2 at qbittorrent.n62.casa. Labels are qBittorrent tags —
movies-main, tv-kids and so on — enough to find things in qBittorrent's own
UI, and distinct from Radarr's existing radarr/sonarr labels so both stacks
can run side by side. Tags rather than a category because a category also
governs the save path, and arr owns that.
torrents/add answers Ok. and nothing else: no hash, no name, and no signal
that the torrent was already there. arr therefore derives the v1 infohash from
the magnet or the .torrent bytes before the call, and looks the torrent up by
it. That is also what makes an add idempotent across a restart (§8) — the same
release resolves to the same hash, and a second add is recognised as the
duplicate it is.
7.2 Hardlink
Verified: Radarr's only media bind is /mnt/media-v2:/mnt/media, one mount,
one dataset, and the existing arrs hardlink-import today. So link() succeeds.
Implementation still falls back to copy on EXDEV rather than trusting that
forever — no configuration flag, no way to get it wrong.
Hardlinking gives the library file and the seeding file independent names at zero extra disk. Copy would double storage for the entire seeding window, which at 4K is 40-80 GB per title.
7.3 Two lifecycles, deliberately decoupled
The torrent and the library entry are separate state machines.
Seeding obligation is per tracker, configured locally because Prowlarr does
not expose tracker rules — ratio and min_seed_time. Set ratioLimit and
inactiveSeedingTimeLimit on the torrent at add time, with
shareLimitAction set to stop rather than delete, and let qBittorrent enforce
them. A reaper deletes torrents qBittorrent stopped on a limit arr set.
Stopped-on-a-limit, not merely stopped: a torrent the operator paused by hand looks identical in the state field alone, and the reaper deletes data. So the ratio and idle limits are checked against the torrent's own counters, and a torrent stopped by a global limit reads as still seeding — the safe direction to be wrong in.
Consequently a hard-failed release is blacklisted and never imported, but its torrent keeps seeding until the obligation clears. Nothing is deleted early to satisfy the library.
7.4 Layout
/mnt/media/movies/main/Dune Part Two (2024) [tmdbid-693134]/
Dune Part Two (2024) [tmdbid-693134] - [2160p][WEB-DL][HDR10].mkv
/mnt/media/tv/kids/Bluey (2018) [tmdbid-82728]/
Season 01/
Bluey (2018) - S01E02 - Hospital [1080p][WEB-DL][pt-PT].mkv
Media kind first, hard audience boundary second, people nowhere.
- Provider ID in the folder name turns Jellyfin matching from fuzzy string guessing into exact lookup.
- Attribute tags come from
ffprobe, not from the release name, so they are true. The filename doubles as an audit surface:lsshows which files are DV or which of the kids' files are still English-only, with the app stopped. - One folder per title even for a single file, so sidecar subtitles and artwork stay contained and deletes are atomic.
- Release group is deliberately absent. It is not a selection criterion and it makes filenames long enough to break a terminal.
Changing a title's root relocates its title folder into the new root; roots are assumed to share one filesystem, so the move is a rename, never a copy. Changing a root's path is the same move over every title under it, and it is all or nothing: one folder that cannot move puts back the ones that already did and leaves the root's path alone, so the stored path always describes the disk.
During transition, write into the existing roots so Jellyfin needs no reconfiguration and new content appears immediately. Radarr will not touch a folder it has no record of.
7.5 Jellyfin
On successful import, call Jellyfin's refresh endpoint for the affected library. Its filesystem watcher misses things. One HTTP call at the end of import.
8. The reconcile loop
There is no job queue.
The database rows are the work list. A wanted episode with no file is a pending grab. A torrent past its seeding rule is a pending delete. A downloaded file not yet probed is a pending import. Every tick, compare desired state to actual state and act on the gap.
This is idempotent and crash-safe by construction. Kill the process mid-grab and the next tick recomputes the same gap and continues. A job table would need retry counts, dead-lettering and reconciliation against qBittorrent anyway, because qBittorrent is an external system that changes underneath the app.
Transient state — a search in flight, download progress — is in memory and rebuilt from qBittorrent on startup. Where that state is ever seen is §9.8: inline on the row that owns the item, never persisted.
Ticks are staggered: reconcile every 30 s, RSS every 10 min, metadata refresh daily, reaper every 5 min.
9. API and UI
9.1 API-first
The HTTP API is the product surface; the web UI is a client of it with no privileged path. OpenAPI spec generated from handler annotations, served alongside the app, and used to generate the TypeScript client.
Own schema first. Radarr-compatible endpoints are a later, separate, lower-priority crate (§9.4), not the primary shape.
9.2 Unified search
One box. Two grouped result sets: in library first, on TMDB below. Both sets are titles — series and movies. Episode rows never appear, under any query, and there are no season rows. Enter on a TMDB result opens the add flow with root and policy pre-filled.
The same box accepts a raw TMDB or IMDb ID, and a pasted magnet or .torrent,
which skips to the manual-grab flow.
There is never a moment where the user has to know whether they are searching or adding.
Result rows are enriched: a poster thumbnail and a rating, under the same rules
as title detail (§9.6) — images hotlinked from path fragments, the rating being
TMDB's vote_average with its vote_count. A trailer chip renders per row and
resolves only when clicked (§9.6).
9.3 Manual search results
Radarr's manual search is unusable because the raw release name is the dominant column, pushing everything that actually decides the choice off-screen.
Invert it. The policy engine has already classified every candidate:
- Three buckets.
eligibleshown by default, sorted by score.waived(fails a soft rule, grabbable with one click that writes an override) andrejectedcollapse to a count, expandable. - Columns are parsed attributes as chips — score, resolution, source, HDR, audio languages, size, seeders. Fixed width, no horizontal scroll.
- Release name is secondary, truncated, full string on expand. It is evidence for when you disagree with the parse, not the primary key.
- Every rejected row names the rule that killed it, so over-strict filters are visible without reading names.
Pre-grab, HDR and audio chips are best-effort from the name (§5.6).
9.4 Jellyseerr compatibility
Jellyseerr stays. It already does Jellyfin user auth, discovery and request approval — none of which is the problem being solved, and rebuilding it doubles the project.
A thin arr-compat crate exposes the slice Jellyseerr actually calls:
GET /api/v3/system/status
GET /api/v3/rootfolder → the real roots
GET /api/v3/qualityprofile → one fake profile per root
GET /api/v3/tag
GET /api/v3/movie existence check
POST /api/v3/movie add
GET /api/v3/movie/lookup
plus the series equivalents. Thin mappings onto the real domain, isolated in
one crate that never leaks into arr-core.
Lower priority than everything in §5-8.
9.5 Notifications
ntfy, one topic per Owner. Three events only — Radarr's failure mode is
notifying on everything and being muted within a week.
- Imported → to the title's owners. The only good-news notification.
- Needs a decision → to the operator alone. Entered the no-PT-source queue, or the needs-a-decision queue (§5.7).
- Broken → to the operator alone. Prowlarr, qBittorrent or TMDB unreachable, disk full.
Not notified: grabs, searches, downloads starting or finishing, soft fails.
9.6 Title detail
One detail surface per kind: /movies/{id} and /series/{id} (#129). For a
movie the release deck becomes a section of the page, and /movies/{id}/releases
keeps resolving; series keeps its season-and-episode shape, seasons ordered
newest-first within the series and episodes newest-first within each season —
the seasons you are deciding about now are at the top, not after nine rows of
back catalogue.
TMDB is the only metadata source. The rating shown anywhere is TMDB's
vote_average with its vote_count, rendered as amber stars (--signal-warn
— the same state colour a wanted chip carries). No OMDb, no IMDb or Rotten
Tomatoes scores — each would need a second upstream, a second key and a second
thing that can be down.
Images are hotlinked from image.tmdb.org. The API returns TMDB path
fragments, never URLs; the browser composes the URL and chooses the size. No
image proxy and no image cache in the service.
Rich detail is not persisted. It is served through arr-meta's 24-hour
response cache, which lives on disk in a directory alongside the database
(§10), bounded by both age and total size (#158) — not the database itself,
and not an image cache; images stay hotlinked as above. The single exception
is poster_path, backdrop_path
and vote_average, stored on movies and series and written by the daily
metadata refresh (§8), so library views render without a TMDB call.
External links are TMDB always, IMDb for movies, TVDB for series — all from ids the app already holds — plus a Rotten Tomatoes search link, which is a query URL, not a resolved title page.
Trailers resolve on click. TMDB's search responses carry no videos, so a trailer key costs a detail call. Rendering one chip per title and resolving the one clicked keeps that cost at one call, and the 24h cache makes a repeat free.
Library view. A poster grid by default with a list toggle; the list keeps the derived-status columns §4.2 built it around.
Out of scope here, because they are the adjacent scope most likely to creep: watch providers, recommendations or similar titles, collections, person pages inside the app, review text.
9.7 The shell
The library is the homepage. / renders the library view described in
§9.6. It is what the operator opens the app to look at, so it is what the app
opens on.
The signal chain is a settings section. The four upstreams — tmdb,
prowlarr, arr, qbittorrent — and their lamps live inside /settings, and have
no route of their own. Per-upstream health is something you check when
something is wrong, not a homepage.
The master lamp stays on the rail, so the at-a-glance read that the chain is healthy survives the move and no navigation is needed to get it.
The wordmark is a link home. arr on the rail navigates to /.
The rail composition is fixed. From left to right it carries the master lamp and wordmark, a warning mark only while queues need attention, the centred search, the version readout and the settings cog.
9.8 Downloads on the row
§8 keeps download progress as transient state; nothing in §9 so far says where it is seen. Without a rule here, a season pack downloading and one that never started look identical.
Progress is inline, on the row that owns the item — a movie row, a season row for its pack, an episode row. There is no downloads page and no count in the rail: a download is an attribute of the thing being downloaded, not a place to visit.
Alongside progress the same row carries the other states a torrent can be in: seeding under §7.3's obligation, stalled, errored.
A torrent arr did not grab is never shown. §2 already rules the service knows only what it put on disk; qBittorrent's own UI lists the rest.
Nothing is persisted — no progress column, no new table. The snapshot comes from qBittorrent and dies with the process (§8), refreshed by the UI's normal polling cadence at roughly 15 s rather than SSE.
At phone width only active-grab progress survives; seeding and stalled shed with the row's other secondary chips.
Out of scope here: any change to what the reconcile loop does, and any new state machine — §7.3's two lifecycles stay as they are.
10. Persistence
SQLite via sqlx, compile-time-checked queries, migrations in arr-db.
A few thousand rows, single writer. Postgres would buy nothing and cost a service.
Policy lives in the database, not a config file — size targets and DV rules get tuned by hand during testing and a restart-to-reload loop gets old immediately. Only bootstrap settings (bind address, Prowlarr URL, qBittorrent URL and login, TMDB key, media root, TMDB response cache directory) come from config/env. The qBittorrent password is a secret, so it is env-only and has no config-file field.
Backup is sqlite3 .backup on a timer.
11. Crate layout
Cargo workspace, members = ["crates/*"], versions pinned once in
[workspace.dependencies], following ~/tea/maestro.
arr-core domain types, policy engine, scoring no IO, no heavy deps
arr-parse release name parsing no IO
arr-meta TMDB client
arr-indexer Torznab via Prowlarr
arr-dl qBittorrent WebUI API
arr-probe ffprobe wrapper
arr-subs subtitle providers, translation, sync
arr-db sqlx + migrations
arr-api axum + OpenAPI
arr-compat Radarr/Sonarr v3 shim for Jellyseerr
arr-daemon reconcile loop, wires everything
arr-e2e cross-process integration tests
web/ Vite + TypeScript SPA, embedded via include_dir
The rule: arr-core and arr-parse hold the logic worth testing constantly and
must not depend on axum, sqlx or reqwest. Everything expensive is downstream of
them.
Nothing is generic over media kind. Movies are built concretely, then TV concretely, and shared machinery is extracted only once both exist. An abstraction derived from one example fits one example.
12. CI
Lints declared once in the root manifest:
[workspace.lints.rust]
unused_crate_dependencies = "warn"
missing_debug_implementations = "warn"
[workspace.lints.clippy]
pedantic = { level = "warn", priority = -1 }
unwrap_used = "warn"
Member crates carry only [lints] workspace = true.
Per-push gate, target under 5 minutes:
| Step | Tool |
|---|---|
| format | cargo fmt --check |
| lint | cargo clippy --all-targets -- -D warnings |
| unused deps | cargo machete (stable; cargo udeps needs nightly and a full rebuild) |
| test | cargo nextest run |
| frontend | biome ci web/ and tsc -b --noEmit |
Off the gate, scheduled: cargo deny for advisories and licenses.
Coverage, if ever, likewise. Neither blocks a push.
Total wall-clock on the self-hosted Gitea runner is an explicit constraint. The
lever is caching, not step selection: cache ~/.cargo/registry, ~/.cargo/git
and target/, keyed on Cargo.lock plus rust-toolchain.toml. Never build
--release in the gate.
End-to-end tests, only at the seams that actually break:
- Prowlarr and TMDB —
wiremockwith recorded real responses as fixtures. Never live: trackers rate-limit, and it would leak credentials into CI. - qBittorrent — a real container, started with a seeded config because the
image prints a random WebUI password per boot. Its API semantics are the most
likely source of surprise:
torrents/addreports nothing back, and required parameters have changed between major versions. ffprobe— tiny committed clips, a few KB each. The Jellyfin LXC already has the right fixtures at/srv/jellyfin-test, including a real DV Profile 5 clip. That is the test proving the policy engine rejects Profile 5 and accepts 8.1 — the subtlest rule in the system.
E2E runs on main and on pull requests touching those crates, not every push.
13. Build order
Each phase ends at something usable end to end. No phase is a refactor of the previous one.
- Skeleton — workspace, CI gate, config, SQLite migrations, health endpoint, embedded empty SPA.
- Parse and score —
arr-parseandarr-coreagainst fixture release names. Pure, fast, heavily tested. No network. - Read-only sourcing — TMDB lookup, Prowlarr enumeration and search, classified results over the API. Still grabs nothing.
- Movies, end to end — add a movie, grab, download,
ffprobe, hardlink, rename, Jellyfin refresh, seeding reaper. The first real cutover test is one movie Radarr does not know about. - UI — unified search, manual search buckets, library views, the queues.
- TV — seasons, episodes,
auto_track, per-episode versus season-pack grabbing, derived status. - Owners and notifications — tags, per-person ntfy topics, filtered views.
- Jellyseerr compat —
arr-compat. - Subtitles — replaces Bazarr. See §15.
Movies before TV because TV adds season packs, air-date calendars and per-episode state on top of an otherwise identical pipeline. Doing it second means that pipeline is already proven.
14. Open questions
- Remux playback. Whether a 4K remux streams cleanly to the Shield is untested. The usual failure is audio, not bitrate: BluRay remuxes carry TrueHD/DTS-HD MA, and if the Shield cannot bitstream that downstream, Jellyfin transcodes audio and playback stutters while video direct-plays. Test before fixing the 4K size ceiling in §5.5. The outcome changes the conclusion from "remuxes are bad" to "remux audio needs a downmix".
- Size band numbers in §5.5 are placeholders pending that test.
- Season-pack re-grab. When an airing season completes, the episodes are already present individually. Nothing re-grabs the pack, and per §6.2 the RSS lane obeys the same guard — it skips packs for any season with episodes on disk. Whether a re-grab is ever wanted remains unresolved and deliberately deferred.
15. Subtitles
Replaces Bazarr. Phase 9 in §13; the arr-subs crate in §11.
Wanted set. Global, not per root. Two languages are separately wanted for every media file: Portuguese — pt-PT preferred, pt-BR accepted — and English. A file is satisfied for a language when a subtitle in it exists, embedded or as a sidecar. Satisfying a language is a statement about viewing it, and not about having text in it — the distinction matters where an image track is all there is, and Translation below is where it bites. This is deliberately unlike §5.2's audio rules, which attach to a root: subtitles carry no blacklist there and none here. pt-BR subtitles are always fine.
Embedded tracks. An embedded subtitle track satisfies its language.
Text-format tracks (subrip, ass, mov_text) are additionally extracted to
a sidecar SRT, because an extracted track is a legal translation source.
Image-format tracks (PGS on BluRay, VobSub on DVD) carry bitmaps, not text:
they satisfy viewing but can never feed a translator, and arr does not OCR
them. arr-probe already reports subtitle tracks with resolved languages; the
format is the new fact it must carry.
Cue skeletons. What an image track does have is exact timings, because
its packet timestamps are the disc's own cue structure. At import — while the
file is being read and hardlinked anyway — arr derives a cue skeleton from
each non-forced image track: ffprobe -show_packets on the track, packets
paired show-to-clear, no pixel read. The skeleton has timings and no text, and
that is enough, because alass matches on interval structure rather than on
words. A sparse skeleton is still a strong reference; a downloaded subtitle and
a retail disc's track never carry the same cues anyway.
The pairing is the whole of it, so it is validated before it is trusted: an even packet count, every implied duration plausible, and a cue density that fits the runtime. PGS permits several composition segments per subtitle, and a track built that way pairs into something that is quietly half a second out — worse than no skeleton. A track that fails validation gets none, and alignment falls back to the video.
Providers. OpenSubtitles.com, behind one trait.
Ranking. A moviehash match wins outright. Then an exact release-name
match, then same release group or same source, then uploader rating and
download count as tiebreakers. Every rejected candidate names the rule that
killed it, so the manual view described in §9.3 works unchanged for subtitles.
Forced and SDH. A forced track covers only foreign-language lines and on-screen signs; it never satisfies a want and arr never goes looking for one. An SDH track is complete and satisfies, ranked below a plain subtitle.
Neither gets its own sidecar name. A language is satisfied by exactly one
sidecar, so <video>.<lang>.srt needs no segment distinguishing forced from
plain from SDH — where a plain subtitle exists for a language, the forced one
is ignored rather than kept beside it. The database says the same thing: one
sidecar row per (media file, language).
Translation. When no provider has a wanted language, arr translates immediately — there is no waiting window. The source is an existing subtitle: a downloaded one, or one extracted from a text-format embedded track. Being able to translate from an embedded track is a deliberate improvement on Bazarr, which cannot.
Fetching a source. A release whose only subtitle is an image track has no text source and never will: the track satisfies its language, so nothing is ever fetched in it, and the track itself cannot be translated from. That is a deadlock, and it is broken by a narrow carve-out — when a wanted language needs translating and no text source exists, arr fetches a text subtitle in the image track's language even though that language reads as satisfied. The fetch obtains a source; it settles no want of its own, and the no-upgrade rule below does not apply to it. Where that track has a cue skeleton, the fetched subtitle is aligned against the skeleton rather than the video, which puts it on the disc's own timings instead of on a heuristic read of the audio. OCR remains a non-goal: this reaches the same place with real text.
Translation backends. Pluggable, each behind its own cargo feature: an
OpenAI-compatible HTTP endpoint, DeepL, Google Translate, and a generic remote
command driven by a configured template (ssh box claude -p is one instance
of that template, not a backend of its own). The engine in use is a database
setting, so switching does not need a rebuild when the feature is compiled in.
Subtitles are sent in batches of cues; a reply whose cue count or numbering
does not match the batch is rejected. Timing data never leaves arr.
No upgrade loop. Once a language is satisfied — by a machine translation too — arr stops working on it. A real subtitle appearing later does not replace anything. Replacement is a manual action from the UI. This is §5.4's rule applied to subtitles.
The one exception is the source fetch above. It is not an upgrade: the language it downloads is already satisfied and stays satisfied by the same track it was before, and what the download settles is a different language's gap. Nothing is replaced, so nothing about the rule changes.
On disk. Sidecars live next to the video inside the §7.4 title folder,
named <video basename>.<lang>.srt, e.g.
… - [2160p][WEB-DL][HDR10].pt-PT.srt. A machine translation carries an extra
.mt segment: … [HDR10].pt-PT.mt.srt. That keeps §7.4's audit-by-ls
property — with the app stopped, the filename says which subtitles are
machine-made. Folder-level delete stays atomic because sidecars are inside the
folder.
.mt is the only optional segment. One language, one sidecar: a second
subtitle for a language arr already has is refused, and replacing one is the
manual delete-then-fetch §15's no-upgrade rule already describes.
Sync. alass runs on every fetched and every translated subtitle. It is a
single small binary invoked like ffprobe, so it costs nothing at rest. It
reports no confidence value, so its output is accepted unless it is
implausible — a shift beyond 60 seconds, or cues lost — in which case the
unsynced original is kept and the file is flagged.
The reference is the video, except where a cue skeleton exists, and then it is the skeleton. A subtitle a skeleton accepted is already on disc-exact timings, and translation copies those timings over untouched, so the pass that would otherwise run on the translation is skipped: a second alignment has nothing left to find and can only move them off.
Configuration. Provider credentials and translator API keys are bootstrap
config or environment, per §10 — a secret never becomes a database row. Wanted
languages, chosen engine, per-provider enable and the daily budgets are
database rows edited from /settings without a restart.
So is everything needed to point the OpenAI-compatible backend somewhere else:
its base URL and its model name are database rows too, not bootstrap
config. That backend is not "OpenAI" — it is any endpoint speaking that shape,
llama.cpp and a local gateway included, and which one is in use is a thing to
try and change, not a property of the deployment fixed at start-up. Only the
API key stays in the environment, and an endpoint that needs no key is a valid
configuration.
Budgets. A token bucket per provider and per translator, with a configured daily allowance. The reconcile loop spends it newest-import-first, so enabling this on an existing library drains the backlog over days instead of hitting every rate limit at once. Being at the cap is a visible queue state, not an error.
Notifications. No new event classes. §9.5 stands: subtitle fetches never notify, and a provider or translator being unreachable folds into the existing "Broken" message to the operator alone.
Non-goals, stated here so they do not creep back: OCR of image-based tracks, transcribing audio when no subtitle exists anywhere, adopting subtitle files already on disk that arr did not write (§2 already says the service knows only what it put there — which does mean arr may fetch a second copy alongside one Bazarr left), and any background loop that upgrades a subtitle in place.