HomelabOS service validation: batches of six¶
Date: 2026-09-09
Source baseline: f951a440e390b14afee1c90ccbcc612e5d60bd7b, feat/service-batch
Status: Batch 1 completed on 2026-09-10. FreshRSS, Linkwarden, Audiobookshelf, Homarr, Dozzle and Uptime Kuma completed their native ARM64 fixture, resource, Authentik, duplicate-instance, lifecycle and isolated-restore checks, including full measured 30-minute paired workloads. The six service guides and consolidated report record their setup, sizing and logout limits. All tested application pairs are stopped with data retained; Dozzle's 20 log fixtures were removed. The shared test infrastructure remains available.
Uptime Kuma finished its paired workload with 37,400 checks and no failures,
347.1 MiB combined sampled peak RAM and 14.82% mean CPU. Two source updates
per instance, stop isolation, Docker container restarts, isolated file/SQLite
verification and fresh post-lifecycle authentication passed. Original cold-start
measurements for Homarr and Dozzle's primaries and both Kuma instances remain
incomplete after harness recovery. Upgrades and full-host/Restic recovery remain
untested. The other 222 catalog entries remain scheduled or in progress. Batch 2
is running: all six images passed native ARM64 admission; Wallos passed primary
startup, settled idle and two active trials (1,145 checks, no failures). Both Wallos
instances passed native SSO/access denial, local recovery and mutation checks;
the duplicate's original startup capture remains incomplete after a routing 404.
Reciprocal API-key/session isolation and the full paired workload passed: 3,439
checks, no failures, 48.64 MiB combined peak RAM and unchanged canonical data.
Wallos also passed two source update/restart cycles per instance, stop isolation,
container restarts, stopped-file/SQLite extraction verification and fresh final
authentication. Both instances are stopped with data retained. Extracted copies
were not booted; its original secondary startup measurement remains incomplete.
Memos primary startup and its 50-memo/ten-comment/five-attachment fixture passed;
native provider setup, its persisted administrator link, fresh Authentik
ADMIN/USER logins, policy/registration denial and local-password recovery
passed. Native logout revoked the refresh session and cleared browser
authentication; an issued access JWT can remain valid for up to 15 minutes.
Global Authentik logout was not tested. Subsequent API/browser and canonical
fixture checks passed. Its settled ten-minute idle window measured 50.28 MiB
median RAM, 51.27 MiB peak RAM and 0.011932% mean CPU, with canonical data and
app/core runtime preserved. Two fresh ten-minute active trials and one-minute
recovery captures passed: 5,892 operations, zero failures, 72.93/72.38 MiB peak
app RAM and 4.17%/4.25% mean app CPU. Both trials preserved canonical data and
app/core runtime. The original secondary startup also passed: administrator ready
in 5.590 seconds, full fixtures in 11.468 seconds and 22.88 MiB startup peak RAM.
Both instances passed native SSO/denial/recovery/logout, API/browser and canonical
snapshot checks, with distinct original persisted signing keys. Both roles passed
reciprocal access-token, refresh-cookie and browser-storage isolation, with own
sessions, canonical data and app/core runtime preserved. The 30-minute paired
workload plus one-minute recovery passed 16,370 operations with zero failures.
Combined app RAM was 124.51 MiB median / 130.29 MiB p95 / 136.25 MiB peak, with
7.29% mean CPU of one core. Both snapshots and all six container configurations
and identities were preserved. The secondary-only note edit and exact restore
passed, with the primary usable and unchanged during the edit, both final API
checks/snapshots, and complete database/attachment/configuration preservation.
Two source deploy/restart cycles per instance, both direct container restarts and
secondary-stop isolation passed, with fresh recovery/SSO and full data,
configuration and peer/core preservation. Both corrected stopped-service backups
and fresh recovery/SSO passed, followed by final denial/logout and retained-data
shutdown with exact shared-core preservation. The first primary backup remains a
recorded reader-sidecar harness failure; extracted copies were not booted and
upgrades/full-host recovery remain untested. Memos's provisional RAM allowance is
256 MiB per instance for these fixtures. HomeBox's disposable migration,
attachment-storage and Docker SQLite schema checks passed. Its two isolated
Authentik providers, access groups and credentials, strict discovery, source
fixtures and overrides are prepared, and the preceding quiet shared-core window
passed. The primary now deploys on ARM64 and initializes OIDC; its initial empty
census passed and registration created the expected owner/group and seed data.
The browser form-transition assertion then failed after HTTP 204. Original pending
evidence and the interrupted capture remain retained. A later continuation
verified native password and API login to the existing owner, with all four
session tokens independently matched and all other database rows unchanged.
Its snapshot then failed on a missing maintenance status filter. A subsequent
session-only continuation completed the read-only snapshot using the existing API
session, but its seeded census failed parsing an empty API asset ID as an integer.
All typed database rows remained unchanged, with no new logins, fixture writes or
API keys. Earlier failures are retained. Full receipt reconciliation and the remaining
fixture operations have now completed: independent inspection verifies 20 items,
11 locations, eight tags, two types, two attachments, one key and four session tokens.
A final read-only verification corrected timestamp precision in historical test
evidence without replaying application operations. The original container and all
journals are preserved. The five-minute settle and ten-minute idle capture passed:
42.57 MiB median / 44.08 MiB peak app RAM, with 0.308378% mean CPU of one core.
Full database rows, journals and runtime remained unchanged. Shared-core usage is
reported separately. Neither these idle figures nor the 89.54 MiB partial startup
peak establish a sizing allowance. Native SSO, complete workload measurements and duplicate reliability
remain pending.
Earlier failed attempts remain separate evidence. See the
pilot report.
Inventory: All services and batch assignments
Goal and scope¶
For each service, establish what works, what it costs to run, whether Authentik provides a working login, whether two independent instances work together, and which setup steps or platform restrictions matter. Publish evidence and useful operator instructions in the service documentation.
The default scope is all 229 tracked service definitions, prioritizing the 58 added roles and eight existing roles relevant to the recent changes. The inventory assigns every service exactly once: 38 batches of six and one final batch of one. A batch is six service families, including their required databases, workers and other dependencies. There is no three-container ceiling.
The user authorized increased parallelism on 2026-09-11. Run independent image, startup and functional checks concurrently, including overlapping batches of six when their combined caps fit the available resources. Use separate owned projects, volumes, source overlays and browser contexts. Coordinate shared deployment and Authentik changes, and reserve measured workload windows so other test traffic cannot distort them. Offline preparation can continue during those windows.
Quiescent static servers may share one idle window, with separate per-container results and a separately reported common core; label this as simultaneous idle. Active and primary/duplicate measurements retain their explicit workload scope. Stop completed test stacks while retaining their data and evidence. Concurrency limits and container caps are test controls, not measured sizing recommendations.
The JSON retains historical inventory/source metadata and now includes current
runtime fields summarizing the linked test evidence. Untested outcomes remain
not_run; image admission alone does not change runtime results.
What is already documented¶
Source inspection on 2026-09-09 found:
| Evidence | Existing coverage | What remains |
|---|---|---|
| Resource footprint section and tracker measurement | 21 of 58 added services; the two lists agree | 37 added services and the older catalog lack this standard section |
| Startup RAM peak | 9 tracker records | Consistent startup sampling for every tested stack |
| Time to first response | 14 tracker records | Measure actual usable readiness, separately from HTTP response |
| SSO section | 48 of 58 added-service docs | Verify the declared protocol, edition, configuration and real login |
| Duplicate-instance wording | 40 of 58 added-service docs | Prove simultaneous operation, data isolation, persistence and separate SSO identities |
| Automated regression suite | 48 offline tests passed before the last push | Live app functionality, login, workload and recovery are separate evidence |
Existing figures are useful historical idle observations, not active-load sizing
or proof of current-version behavior. The tracker has no standard CPU/workload
series, architecture/digest fields or separate SSO/isolation outcomes. Its
mr-open status describes contribution state, not deployment reliability.
There are two normalized WriteFreely tracker rows: Write-Freely contains the
completed measurements and WriteFreely is marked duplicate/skip. Preserve both
rows and join by explicit role identity; do not overwrite one during an import.
Local environment and isolation¶
Observed when the user authorized execution on 2026-09-09:
- Mac: ARM, 64 GiB RAM, 12 CPUs.
- Docker Desktop 4.90.0: 23,296 MiB configured RAM / 22.2 GiB reported by the
engine, 12 CPUs,
aarch64, 256 GiB virtual disk; Engine 29.7.2 and outer Compose v5.4.0. Record the inner Engine/Compose versions separately. - Existing
hlos-testserveris stopped. Other workloads include Grownetics, Fieldwalk and Valhalla; they must not be reset, stopped or used as fixtures. - The initial planning snapshot had 7.65 GiB VM RAM. After a manual update, stale backend processes referencing a deleted update-staging directory prevented startup. Relaunching the installed app restored the engine; RAM was lowered from 32 to 16 GiB at the user's request and the disk was preserved. The user subsequently increased memory before starting the campaign. This was an environment repair, not service-validation evidence. Recheck available capacity and emulation support when starting the campaign.
Use a new, clearly named test target such as hlos-batch-<run-id> based on the
existing Ubuntu/systemd/Docker-in-Docker approach. Give it its own image-cache,
containerd and application-data volumes. Preserve the old test target and data.
Keep app data on Linux volumes inside the test environment for consistency.
Run the real HomelabOS Ansible/systemd deployment path from a source-only export of the selected commit. Generate a fresh inventory, settings and synthetic secrets inside that export. Do not load the working checkout's real settings or vault. Record the source SHA and any test-only overrides with every result.
All cleanup must resolve containers, volumes, networks and directories from the
run's ownership manifest. Never use host-wide prune/reset commands. The existing
reset_testenv.sh is hardcoded to hlos-testserver and scans all its service
directories; it must be adapted or replaced before use with the new target.
Before admitting concurrent work, measure VM available memory, Docker disk space and background/core usage. The expanded parallel campaign reserves at least 4 GiB of VM available memory, approximately two CPU cores and 20 GiB of Docker disk space. Account for all newly admitted app/dependency caps and test clients together, rather than granting each batch the same free capacity. Unknown heavy stacks start singly. Stop the current test on memory pressure, OOM or unrelated-workload degradation and record an environment/capacity result. Reassess the budget from observed usage; do not automatically increase Docker Desktop's allocation or restart it. This supersedes the initial 1.5 GiB reserve.
Docker-in-Docker shares the Desktop Linux kernel. Network appliance, VPN, device, GPU and host-management tests may require a separate Linux VM or physical host. A successful nested-container test does not certify those host integrations.
Phase 0: establish trustworthy infrastructure and measurement¶
Complete this before declaring failures in the first batch.
- [x] Capture source revision, hardware/VM allocation, inner and outer Docker versions, free disk, background activity, and exact image digests/platforms.
- [x] Create the isolated target, source export, test accounts, ownership manifest and persistent evidence directory.
- [x] Deploy Traefik and Authentik, including all required Authentik dependencies. Deployment/startup and a separate quiet-core idle window are complete. The settled ten-minute idle window measured 1,321.6 MiB median / 1,323.2 MiB p95 RAM and 2.01% mean CPU, with no restarts or OOM events.
- [x] Establish one canonical HTTPS origin for each app and for Authentik that works from both the dedicated test browser and application containers. The pilot uses isolated ARM64 Chromium on Docker with its own trusted CA store.
- [x] Validate certificate trust, discovery, JWKS, authorization, token exchange, callback and application session with a dedicated test client. Confirm a real browser login in a pilot app before using the infrastructure for other apps.
- [x] Enable the embedded proxy outpost when required and verify actual browser authentication. Record a settled core baseline before proxy setup and concurrent four-container core measurements during Dozzle and Uptime Kuma. The embedded outpost has no separate container; observed core differences do not isolate its causal resource cost. A provider object alone is not a login.
- [ ] Test an LDAP outpost when a later service requires it. An isolated causal estimate of proxy-outpost overhead has not been established by this pilot.
- [x] Correct the measurement/runner limitations below and verify their behavior on known fixtures before taking resource measurements.
Local SSO routing¶
Preferred browser names are <instance>.127-0-0-1.sslip.io, with standard HTTPS
port 443 exposed only on Mac loopback. Use a separate SSH test port. Check port
availability immediately before startup; the planning inspection found no
listeners on 80, 443 or 2223.
Inside the app network, those same names must resolve to the test Traefik endpoint, not container loopback. Add explicit network aliases or test-only DNS records for Authentik and both instances. Ensure every required worker/back-end can use that resolver. Authentik's canonical issuer must remain the same from the browser and app side.
Use a dedicated development CA/certificate trusted by the test browser and app trust stores. Record any required CA installation steps. Do not treat disabled TLS verification as a successful secure SSO test. If standard ports are busy, use one explicitly configured alternative port end to end, including internal reachability, issuer, external URLs and callbacks; prove that topology first.
The old HTTP-only curl -H Host ... localhost:8080 check cannot validate this.
In particular, a loopback discovery hostname inside a container may resolve to
the app itself. Authentik documents the separate browser and server-to-server
parts of OIDC and the need for accessible metadata/JWKS/token endpoints.
(Authentik OAuth2/OIDC)
Required improvements to existing helpers¶
homelabos-pipeline/measure.py and the test runbook are starting points, not yet
the final measurement harness. Before execution:
- Make the target container and canonical app URL explicit parameters.
- Select containers using Compose project/service labels and the run manifest,
not
slug_name prefixes. Include auxiliary services and record exited one-shot jobs; do not confusefreshrsswithfreshrss2. - Surface Docker/measurement errors as errors. The current helper returns empty output on failed Docker commands, which can look like a zero-container result.
- Separate first HTTP response from usable readiness. The existing watch loop accepts unexpected codes such as 500 and takes one reading after 20 seconds; it is insufficient for this plan.
- Stream timestamped samples from before first start through readiness, settling, workload and restarts. Save raw samples and calculate aggregates.
- Standardize bytes/MiB/GiB and label legacy
MBfields explicitly. Track image size separately from writable/data storage; account for shared images/layers. - Capture exit codes, health changes, restart counts and OOM events alongside metrics, and export structured, sanitized evidence.
The pilot's new collector and window analyzer now cover these measurement prerequisites, including 44 passing offline tests. The old helper is retained only for historical use. Fixtures cover project selection, Docker errors, one-shot completion, HTTP 500 rejection, unit conversion and lifecycle races.
Repeatable procedure for each service¶
1. Preflight the exact deployment¶
- Read the current role and official documentation for its pinned release/edition. Record the official source URL and deliberate deviations from its deployment.
- Resolve every app/database/worker image to a digest and inspect all required architecture manifests. Floating tags must be pinned to the observed digest for a run; a retest with a new digest is a new result.
- Render primary and duplicate with minimal settings; validate the result using the Compose executable actually installed in the test target. Verify network declarations, host ports, paths, permissions and lifecycle/dependency semantics.
- Classify SSO mode, first-admin steps, external fixtures, paid-edition features, hardware requirements and expected representative operation before startup.
- Preserve required dependencies even if the stack needs more than three containers. A missing ARM database image can block an otherwise ARM-native app.
2. Cold start and first administrator¶
- Pull/cache images first and record pull duration separately. Start sampling before the service unit starts with fresh, owned application data.
- Record migrations, health transitions, first response, and time until the app supports a meaningful operation. Require expected status/content or API output, not any response or a setup redirect.
- Complete first-run setup and prove that the intended test administrator can administer the app. Check a normal user cannot gain those privileges.
- Default startup limit: 20 minutes; declare 45 minutes for known heavy builds/indexers before the run. After a reproducible fault, spend at most 30 minutes diagnosing it before recording a blocker and continuing the batch.
3. Measure idle and representative activity¶
After readiness, allow at least five minutes to settle, then collect ten minutes of idle samples. Exercise a documented, repeatable workload for ten minutes and measure recovery afterward. Repeat the workload once; retain both runs and any variation. If startup is still active, extend settling or mark idle provisional.
| Measurement | Definition |
|---|---|
| Idle RAM/CPU | Median and p95 of summed current service-stack samples, including its DB/workers |
| Startup and active RAM | Maximum observed sampled stack usage; record interval and any OOM |
| CPU | Average, p95, maximum, host CPU count and convention; 100% means one logical CPU in the chosen Linux collector |
| Timing | Pull time, first response, usable readiness, warm restart and cold restart |
| Disk | Unique image logical sizes, isolated-daemon shared-layer accounting, persistent data before/after workload |
| I/O | Network/block bytes and workload duration, with dataset and concurrency |
| Pair cost | Simultaneous primary + duplicate measured together; common Authentik/Traefik cost listed separately |
Use roughly 1-second startup sampling and 2-second steady-state sampling; record actual intervals and missed samples. A sampled maximum is not a guaranteed instantaneous peak. Docker's Linux CLI subtracts cache from displayed memory; state whether metrics use that convention or raw cgroup/API values. (Docker stats)
Example fixture profiles, finalized for each role before execution:
| Service family | Representative operation |
|---|---|
| RSS/bookmarks | Import 10 local fixture feeds / 100 saved pages, refresh and search; two test users |
| Documents/notes | Upload or create 50 synthetic documents, edit, search and download |
| Media | Import a small licensed/synthetic library and play a sample; test direct play and any supported transcoding separately |
| Dashboard/monitoring | Configure 20 synthetic items/targets and observe three polling cycles |
| Developer/database tools | Create a project/database or repository, write/read records, exercise one job |
| File sync/sharing | Upload/download a fixed 100 MiB test corpus and compare checksums |
| Service-specific systems | Declare equivalent synthetic work and required agents/backends; if unavailable, record exactly which capabilities remain untested |
Publish an observed operating range for the stated fixture, plus a clearly labeled provisional headroom estimate. Do not turn one idle number into a production minimum or extrapolate linearly to arbitrary users/data. All required sidecars count. Shared infrastructure is charged once in host-size summaries.
4. Test Authentik end to end¶
Record integration mode separately from its result:
- Native OIDC or SAML: complete browser authorization, callback and actual app session; verify identity and the expected normal/admin permissions.
- LDAP: prove directory-backed login and group/user behavior; do not label it browser SSO when a separate login prompt is required.
- Forward-auth: prove an anonymous request is blocked/redirected and an allowed Authentik user can reach the app. Record whether the app still requires its own credentials. Access protection alone is not native account login.
- No native login/stateless app: record proxy protection as a separate option.
- Paid edition, unsupported release, unknown capability or manual setup: state the exact restriction. A missing integration guide is not evidence of no SSO.
Test an allowed user, a denied user, repeat login, and app logout. Record whether logout also ends the IdP session instead of assuming global logout. Check local administrator/recovery behavior, first-time user creation and existing-account linking where supported. Auth-dependent APIs/mobile clients need a separate check; do not assume browser SSO authenticates those APIs.
For proxy integrations check every enabled alternate route for bypass and verify that client-supplied identity headers cannot impersonate users. Use disposable fixtures and the embedded outpost in the isolated test Authentik instance; preserve its other provider attachments and configure its public HTTPS browser origin for the pinned version. Dozzle's Docker access targets only the inner test daemon. (Authentik forward auth)
For shared OIDC provisioning, rerun deployment: verify no duplicate provider/app, stable per-instance credentials, correct callbacks, and preservation of manually added test settings. Log API outcomes without logging credentials or tokens.
5. Prove duplicate-instance isolation and persistence¶
The required scenario is the user's example: freshrss and freshrss2, with
source_service: freshrss on the latter. Apply the equivalent naming to each role.
freshrss2:
enable: true
source_service: freshrss
Start with that minimal duplicate, then explicitly set its own auth,
https_only, custom domain and service-specific options for the relevant cases.
Those access settings do not inherit automatically in the current deploy tasks.
To claim reliable operation under this test, require:
- Both instances run simultaneously, each with its own URL and administrator.
- Each has distinct writable paths, secrets, DB identity, Compose project, container names where fixed names exist, and scheduled jobs. Shared read-only fixtures are allowed only when recorded.
- Write different sentinel records/files to each. Neither sees the other's private data; each logs in to the correct instance/provider.
- Enable SSO for both when supported. Check separate client IDs, callbacks and identities; changing the second instance's settings must not change the first.
- Apply the same configuration twice; credentials and data stay stable and provisioning does not create duplicate objects.
- Complete two stop/start cycles, including one container recreation without deleting data, followed by another operation and login in each instance.
- Stop the second instance; the primary stays healthy. Restore a backup of the synthetic data into an isolated target and verify records/checksums.
- After standalone evidence is saved, run a paired 30-minute soak and repeat representative operations. Record pair RAM/CPU and any cross-instance effects.
Distinguish works, works_with_documented_configuration, failed,
not_applicable, and blocked_on_this_environment. A host-wide singleton must
state why a duplicate is inappropriate. Render-only success is never sufficient.
Upgrade/migration reliability requires its own version-pair test and remains
not_tested unless actually performed.
6. Record, document, and clean up¶
Save raw evidence before stopping the test's own units. Leave unrelated workloads and the legacy test target untouched. Retain enough owned state to reproduce failures; do not erase failed data automatically.
For each service update its source roles/<slug>/docs.md with these sections:
- Validation: date, source SHA, image digest/version, architecture, test type and explicit pass/fail/untested outcomes.
- Resource footprint: idle, startup and active figures; full stack/container count; workload; measured pair cost; limits of the sizing estimate.
- Authentik: protocol, tested result, automation versus manual steps, callback/issuer requirements, administrator setup and remaining login prompts.
- Multiple instances: a working
source_serviceexample, isolation result, required unique paths/ports/credentials and unsupported features. - Special requirements: CPU architecture, hardware, external services, agents, SMTP, licensing/edition, permissions, migration and recovery notes.
Preserve historical measurements with their dates; label stale/missing data. Do not replace measured values with guessed ones to fill a table.
Use a new run directory under homelabos-pipeline/test-results/<run-id>/ for raw
metrics, sanitized logs, rendered-config hashes, screenshots and per-case JSON.
Keep test passwords/tokens/private certificates out of published evidence.
Retain synthetic secrets only in an ignored private run directory for retesting.
Maintain a consolidated service-validation table in the HomelabOS docs with
independent columns for runtime, resources, SSO, multi-instance and exceptions.
Keep tracker.csv contribution statuses intact; link its notes to validation
results. Use the new inventory for older roles not represented in the tracker.
At each batch boundary: publish six outcome rows, identified fixes, deferred
tests and updated service docs; verify evidence links, rerun affected offline
checks, and commit reviewed durable changes on feat/service-batch. Retest fixed
services at the new SHA. Push using the established branch workflow when the
batch is ready; do not conflate pushing source with deploying a live homelab.
Architecture and failure classification¶
Each test dimension gets its own status: passed, failed, not_run,
blocked, not_applicable, or partial. Attach a reason and evidence:
| Reason | Meaning and next step |
|---|---|
| Native ARM64 | Exact required images and features run natively; eligible for Mac resource measurements |
| AMD64-only | One or more required images lack ARM64; schedule a native AMD64 retest |
| Emulated | Explicit optional compatibility run; label results and exclude them from native sizing estimates |
| Hardware/host integration | Requires devices, GPU, privileged kernel/network behavior or a real host; identify the appropriate retest environment |
| External dependency/edition | Missing test agent, upstream fixture, account or licensed feature; distinguish app startup from that feature |
| Test infrastructure | DNS, certificate trust, inner Docker/systemd or capacity prevented the test; repair or defer, not an app failure |
| Product/role defect | Reproducible against valid prerequisites; save exact trigger and source fix/retest |
Docker Desktop supports multi-platform emulation, but the actual nested test environment must be checked rather than inferred from the host or old notes. Keep emulated performance separate; Docker documents substantial emulation cost for compute-heavy work. (Docker multi-platform support)
Special cases already needing attention:
- FreshRSS 1.29.1: the initial docs claimed Debian/x86_64-only OIDC. The
pilot verified
mod_auth_openidcand complete native login on the pinned ARM64 Debian image; current service docs record that result. Provider setup remains manual because the role does not call the shared provisioner. First-admin HTTP Authentication setup, distinct per-instance OIDC configuration and independent API passwords passed in this fixture scope. (Release OIDC guide, release Dockerfile) - Dozzle/Uptime Kuma: current docs describe forward-auth; record protection separately from app account login. Dozzle actions must touch only disposable inner-daemon containers.
- Nextcloud: known fixed container names/shared Documents path need actual isolation checks; do not pre-label a minimal duplicate as reliable.
- Rocket.Chat, AFFiNE, Wiki.js, WriteFreely, Pingvin Share: audit findings cover DB/image compatibility, extensions or maintenance. Test in disposable stacks; record findings without silently changing production database major versions.
- NetBird, DNS/VPN roles and host utilities: use isolated peers/test networks; never change the Mac's or existing workloads' routing to make a test pass.
- Six known invalid-YAML roles: hubzilla, keycloak, listmonk, mailu, octoprint, turtl fail preflight until repaired. Keep them in the inventory with that result.
- Nonstandard/missing Compose: cockpit uses host apt tasks; matrix, nut and unofficial_ddns lack the conventional role Compose path. Inspect their actual deployment mechanism; absence is not automatically an ARM limitation.
Batch schedule¶
The first 11 batches contain all 58 added roles plus FreshRSS, Outline, Organizr, Nextcloud, Immich, Authentik, Authelia and Pi-hole. Authentik is first established as infrastructure in phase 0; its own full role/duplicate assessment remains in batch 11. Later catalog batches are ordered by slug, subject to preflight deferrals. A blocked service retains its slot and outcome rather than silently disappearing from the report. Retests are separately linked to its original slot.
| Batch | Six services (except the final remainder) | Focus |
|---|---|---|
| 1 | freshrss, linkwarden, audiobookshelf, homarr, dozzle, uptimekuma | Pilot: validate the test process |
| 2 | wallos, memos, homebox, navidrome, ittools, excalidraw | Lightweight and local-data apps |
| 3 | calibreweb, kavita, komga, seerr, prowlarr, changedetection | Libraries and media integrations |
| 4 | forgejo, joplin, papra, filerise, tandoor, kitchenowl | Persistent data and application workflows |
| 5 | directus, nocodb, baserow, kanboard, planka, docmost | Databases and collaborative apps |
| 6 | wordpress, drupal, dokuwiki, wikijs, writefreely, pingvinshare | Publishing and maintenance exceptions |
| 7 | semaphore, termix, nexterm, openobserve, pgadmin, beszel | Operations and endpoint integrations |
| 8 | dashy, arcane, maintainerr, zipline, openwebui, snipeit | Management and optional features |
| 9 | affine, adventurelog, infisical, convex, appsmith, weblate | Heavier stacks and dependency readiness |
| 10 | rocketchat, netbird, outline, organizr, nextcloud, immich | Migration and instance-isolation hotspots |
| 11 | karakeep, seafile, stirlingpdf, authentik, authelia, pihole | Remaining additions and infrastructure roles |
| 12 | actual, adguardhome, airsonic, apache2, archisteamfarm, archivebox | Remaining catalog; preflight before execution |
| 13 | backrest, barcodebuddy, bazarr, beets, bookstack, calibre | Remaining catalog; preflight before execution |
| 14 | chowdown, clawdbot, clickhouse, cockpit, codeserver, codimd | Remaining catalog; preflight before execution |
| 15 | digikam, drone, duckdns, duplicati, elkstack, emby | Remaining catalog; preflight before execution |
| 16 | erpnext, esphome, ethercalc, factorio, fileflows, firefly_iii | Remaining catalog; preflight before execution |
| 17 | folding_at_home, frigate, funkwhale, ghost, gitea, gitlab | Remaining catalog; preflight before execution |
| 18 | gluetun, gomodel, gotify, grafana, graylog, grocy | Remaining catalog; preflight before execution |
| 19 | grownetics, guacamole, healthchecks, heimdall, hermes, hlos_dash | Remaining catalog; preflight before execution |
| 20 | homeassistant, homebridge, homedash, hubzilla, huginn, invidious | Remaining catalog; preflight before execution |
| 21 | invoiceninja, invoiceplane, jackett, jellyfin, jenkins, keycloak | Remaining catalog; preflight before execution |
| 22 | kibitzr, langfuse, langsmith, lazylibrarian, lidarr, listmonk | Remaining catalog; preflight before execution |
| 23 | mailu, massivedecks, matomo, matrix, matterbridge, mayan | Remaining catalog; preflight before execution |
| 24 | mealie, minecraft, minecraftbedrockserver, miniflux, minio, monicahq | Remaining catalog; preflight before execution |
| 25 | mstream, mybb, mylar, n8n, netdata, nodered | Remaining catalog; preflight before execution |
| 26 | ntfy, nut, nzbget, nzbhydra2, octoprint, odoo | Remaining catalog; preflight before execution |
| 27 | ollama, ombi, opencode, opengsd, openldap, openvpn | Remaining catalog; preflight before execution |
| 28 | overseerr, ownphotos, paperless, paperless_ngx, paseo, peertube | Remaining catalog; preflight before execution |
| 29 | photoprism, phpbb, piwigo, pixelfed, pleroma, plex | Remaining catalog; preflight before execution |
| 30 | portainer, postgresql, privatebin, privoxyvpn, prometheus, qbittorrent | Remaining catalog; preflight before execution |
| 31 | quakejs, radarr, readarr, restic, rsshub, sabnzbd | Remaining catalog; preflight before execution |
| 32 | samba, searx, searxng, seat, shinobi, sickchill | Remaining catalog; preflight before execution |
| 33 | simplyshorten, snibox, sonarr, speedtest, speedtest_tracker, statping | Remaining catalog; preflight before execution |
| 34 | sui, syncthing, taisun, tautulli, teedy, thelounge | Remaining catalog; preflight before execution |
| 35 | thespaghettidetective, tick, tiddlywiki, transmission, trilium, tubearchivist | Remaining catalog; preflight before execution |
| 36 | tubearchivist_jf, turtl, ubooquity, unificontroller, unofficial_ddns, vaultwarden | Remaining catalog; preflight before execution |
| 37 | vikunja, wallabag, watchtower, webdavserver, webtrees, webvirtmgr | Remaining catalog; preflight before execution |
| 38 | wekan, workadventure, xfinityusageinfluxdb, xteve, zammad, ztncui | Remaining catalog; preflight before execution |
| 39 | zulip | Remaining catalog; preflight before execution |
Batch 1 acceptance work¶
| Service | Core operation | SSO and isolation focus |
|---|---|---|
| FreshRSS | Import fixture feeds, refresh, read/search, test API credentials | Inspect ARM OIDC module; exact first-admin setup; distinct freshrss/freshrss2 feeds and providers |
| Linkwarden | Save/search a fixture bookmark and retrieve its archived content | Native OIDC, correct origin/callback and readiness of required sidecars |
| Audiobookshelf | Import/play a small licensed or synthetic audio sample | Complete documented app-side setup; library and user isolation |
| Homarr | Create a dashboard and fixture widget, reload after restart | Native OIDC and administrator mapping; separately record absent Docker widgets |
| Dozzle | Read logs from disposable fixture containers | Authentik proxy protection, anonymous/alternate-route checks and protected actions |
| Uptime Kuma | Create a local fixture monitor and observe a state change | Proxy protection versus separate app login; monitor/status-page isolation |
Completion criteria and next action¶
The program is complete when every inventory entry has an evidenced outcome for each dimension, current service docs, and a concrete retest reason/environment for every blocked or untested feature. This does not require pretending every service works on ARM, or declaring a failed app reliable after an HTTP response.
Phase 0 and the first six-service pilot are complete within the recorded scope. The current batch is Wallos, Memos, HomeBox, Navidrome, IT Tools and Excalidraw; image admission alone is not a complete service pass. Completed outcomes and explicit limits are recorded in the validation report. The plan describes acceptance criteria; only recorded test evidence establishes an outcome.