Knowledge base
what the system has learned — measured, never asserted
Incident mix × tree coverage (30d)
| Incident signature (30d) | Incidents | Tree |
|---|---|---|
| chronic_foreign_heal:ghostshell-searxng | 31 | covered |
| chronic_foreign_heal:ghostshell-prometheus | 30 | covered |
| chronic_foreign_heal:ghostshell-redis-exporter | 30 | covered |
| chronic_foreign_heal:ghostshell-cadvisor | 30 | covered |
| chronic_foreign_heal:ghostshell-postgres-exporter | 30 | covered |
| chronic_foreign_heal:ghostshell-node-exporter | 29 | covered |
| chronic_foreign_heal:ghostshell-alertmanager | 29 | covered |
| chronic_foreign_heal:docker:zombie_containers | 28 | covered |
| chronic_foreign_heal:ghostshell-luks-exporter | 28 | covered |
| chronic_foreign_heal:docker:no_restart_policy | 25 | covered |
| silent_failure:-:pg_cache_hit_ratio | 23 | no tree |
| silent_failure:-:pg_db_size_bytes | 23 | no tree |
| chronic_foreign_heal:service:redis | 21 | covered |
| chronic_foreign_heal:service:postgres | 20 | covered |
| dashboard_health:ghostshell-api-guard:DASHBOARD_UNREACHABLE | 12 | no tree |
| silent_failure:-:pg_conn_used_pct | 11 | covered |
| posture_brute:172.20.0.14 | 10 | covered |
| chronic_foreign_heal:probe:postgres_pool | 6 | covered |
| chronic_foreign_heal:ghostshell-chromadb | 6 | covered |
| chronic_foreign_heal:ghostshell-dashboard | 6 | covered |
| chronic_foreign_heal:docker:unhealthy_ghostshell-luks-exporter | 4 | covered |
| chronic_foreign_heal:service:chromadb | 4 | covered |
| dashboard_health:ghostshell-db-guard:DASHBOARD_UNREACHABLE | 4 | no tree |
| chronic_foreign_heal:probe:nats_stream | 4 | covered |
| anatomy_new:workflow|v#_heal_audit: 'ghostshell-postgres-exporter' heal success #% over #d (# events) | 2 | no tree |
| anatomy_new:workflow|v#_heal_audit: 'ghostshell-luks-exporter' → 'healed' #× in #d | 2 | no tree |
| anatomy_new:workflow|v#_heal_audit: 'ghostshell-prometheus' → 'healed' #× in #d | 2 | no tree |
| anatomy_new:workflow|v#_heal_audit: 'docker:no_restart_policy' heal success #% over #d (# events) | 2 | covered |
| anatomy_new:workflow|v#_heal_audit: 'ghostshell-luks-exporter' heal success #% over #d (# events) | 2 | no tree |
| chronic_foreign_heal:ghostshell-caddy | 2 | covered |
| anatomy_new:workflow|v#_heal_audit: 'ghostshell-postgres-exporter' → 'healed' #× in #d | 2 | no tree |
| anatomy_new:workflow|v#_heal_audit: 'ghostshell-searxng' → 'healed' #× in #d | 2 | no tree |
| anatomy_new:workflow|v#_heal_audit: 'ghostshell-searxng' → 'escalated' #× in #d | 2 | no tree |
| anatomy_new:workflow|v#_heal_audit: 'ghostshell-cadvisor' heal success #% over #d (# events) | 2 | no tree |
| anatomy_new:workflow|v#_heal_audit: 'ghostshell-node-exporter' → 'healed' #× in #d | 2 | no tree |
| anatomy_new:workflow|v#_heal_audit: 'ghostshell-cadvisor' → 'healed' #× in #d | 2 | no tree |
| chronic_foreign_heal:probe:consumer_bindings | 2 | covered |
| anatomy_new:workflow|v#_heal_audit: 'ghostshell-redis-exporter' → 'healed' #× in #d | 2 | no tree |
| anatomy_new:workflow|v#_heal_audit: 'ghostshell-redis-exporter' heal success #% over #d (# events) | 2 | no tree |
| aging_incident:ddb354c8-f0f6-418f-8727-a1087965d9a6 | 1 | no tree |
| aging_incident:27380fe8-9d59-4643-8243-34e2ec007d20 | 1 | no tree |
| anatomy_new:workflow|v#_heal_audit: 'ghostshell-alertmanager' heal success #% over #d (# events) | 1 | no tree |
| posture_brute:171.231.197.30 | 1 | covered |
| aging_incident:a340028a-05a0-4f49-87f5-28a6a6a68463 | 1 | no tree |
| aging_incident:14c85fc8-369c-4392-9fd6-550bf7ba4c76 | 1 | no tree |
| anatomy_new:workflow|v#_heal_audit: 'ghostshell-luks-exporter' → 'escalated' #× in #d | 1 | no tree |
| aging_incident:dc4b1a77-c701-4171-badf-680bafc8a8fb | 1 | no tree |
| aging_incident:ab016aa1-993d-453a-a2ab-14f10b6a3312 | 1 | no tree |
| anatomy_new:workflow|v#_heal_audit: 'ghostshell-v#-api' → 'escalated' #× in #d | 1 | no tree |
| anatomy_new:workflow|v#_heal_audit: 'ghostshell-node-exporter' → 'escalated' #× in #d | 1 | no tree |
Learned thresholds (latest per metric)
| Metric | Service | Latest value | Previous | % change | Method | Confidence | Computed |
|---|---|---|---|---|---|---|---|
| gs_heal.failed REVIEW | probe:nats_stream | 1.0 | 1.5 | -0.33 | median | — | 2026-07-22 01:14:37 UTC |
| gs_heal.healed | secret:admin_password.txt | 2.0 | 2.5 | -0.2 | median | — | 2026-07-21 21:07:18 UTC |
| pg_conn_used_pct | — | 9.0 | 9.25 | -0.03 | median | — | 2026-09-15 16:11:34 UTC |
Skills & autonomy
| Skill | Type | Target | Risk | Reversibility | Ceiling | Level | Host | Wilson LB | Stats | Promoted |
|---|---|---|---|---|---|---|---|---|---|---|
| analyze_table | fix | postgresql | low | IDEMPOTENT_SAFE | auto | propose (default) | — | — | — | — |
| block_brute_force_ip | fix | network | medium | REVERSIBLE | auto | confirm | ghostshell-host | 0.0 | 0✓ / 0✗ / 0↩ / 0 flaps / 0 sigs | 2026-07-22 10:52:37 UTC |
| bound_long_running_job_steps Guided repair for a cron job repeatedly killed by its own hard timeout (press-review-scrape 2026-07-27..31: TIMEOUT after 21600s four times in five days — ~67 spiders, sequential per country group, and ONE hanging spider (walfnet.com TCP timeouts burning retries) stalled its whole group past the 6h cap; a healthy run is 3.3h). Steps: (1) read the job log backward from the kill to find the sub-task that was running when time expired — the cap is almost never the defect. (2) Bound EACH sub-task, not the job: `timeout --signal=TERM --kill-after=30 <budget> <cmd>` per step, sized ~2-3x its healthy duration (reference: scrape-all.sh SPIDER_TIMEOUT=900, 2026-08-02). (3) Preserve partial results: if the sub-task commits incrementally (scrapy -> DB), fall through to the export/assembly step on timeout instead of aborting. (4) Only then consider raising the job cap — and say the number. Verify: next three runs complete under the cap with per-step timeouts logged. | investigation | generic | low | IDEMPOTENT_SAFE | confirm | propose (default) | — | — | — | — |
| cancel_query | fix | postgresql | low | IRREVERSIBLE_RECOVERABLE | auto | propose (default) | — | — | — | — |
| clean_temp_files | fix | linux | low | IRREVERSIBLE_DESTRUCTIVE | confirm | propose (default) | — | — | — | — |
| fix_backup_false_failure Guided triage for a 'backup failed' alert (SOWKNOW weekly restic 2026-08-02: alert fired, but snapshot 4ee0bfd7 (22.6 GiB) WAS in the repo — restic exited 1 because transient files vanished mid-scan in the LIVE PostgreSQL data dir: 'lstat .../base/<oid>: no such file or directory'). Restic exit codes: 0 clean, 1 = warning (snapshot saved, some sources unreadable), >=3 = fatal. Steps: (1) NEVER trust the alert alone — check the repo first: `restic snapshots --tag <tag> --last 3`; if a fresh snapshot exists, the page was false. (2) Fix the script to treat exit 1 as success-with-warning and only fail on >=3 (reference: sowknow4 scripts/backup.sh + backup-full.sh, 2026-08-02). (3) Reduce the warning class itself: prefer backing up pg_dump output over the live data dir, or exclude transient relation files. (4) Keep the freshness watchdog (alert when newest snapshot > 25h) — that is the check that catches REAL silent backup death. | investigation | linux | low | IDEMPOTENT_SAFE | confirm | propose (default) | — | — | — | — |
| fix_container_dns_resolution Guided repair for container restart loops caused by DNS ('[Errno -3] Temporary failure in name resolution' — Sowknow backend↔vault loop, 2026-07-24). Docker's embedded DNS only resolves names of RUNNING containers ON THE SAME network: a dependency that is down, profile-gated off, or on another compose network makes the client's health check fail and docker's restart policy loops the container — restarting can never fix a name that does not resolve (Ghostshell RCA lesson #30). Steps: (1) from inside the looping container, `getent hosts <dep>` / `nslookup <dep>` to confirm the miss; (2) `docker network inspect <net>` on BOTH containers to compare memberships; (3) `docker ps -a` — is the dependency running at all (Sowknow's vault is profile-gated OFF by design)? (4) The fix is app-side and exactly one of: start/attach the dependency, point the client at the right hostname/network, or make the health check treat an OPTIONAL dependency as optional. Then verify: container stays Up > 10 min, health check 200. | investigation | linux | low | IDEMPOTENT_SAFE | confirm | propose (default) | — | — | — | — |
| fix_duplicate_cron_jobs Guided repair for the same pipeline scheduled in MULTIPLE users' crontabs (Ghostshell press-review 2026-08-02: root AND mamadou both ran press-review-scrape/generate at the same times; mamadou's copies failed on every run with 'open .env: permission denied' — the env file was 600 another user — spamming failure summaries, and any run that DID succeed would have double-sent the report). Steps: (1) enumerate ALL crontabs: `for u in $(cut -d: -f1 /etc/passwd); do crontab -l -u $u; done` plus /etc/cron.d, /etc/cron.daily, systemd timers — duplicates hide across mechanisms. (2) Group by schedule+command; keep exactly ONE owner per pipeline: the user that can read the pipeline's env/secrets files. (3) Move plaintext credentials OUT of crontabs into a 600 env file sourced by the job (`set -a; source .env`) — a password in a crontab is readable by any local user via `crontab -l` on multi-admin hosts and leaks into logs. Watch for space-containing values (Gmail app passwords) — quote them or sourcing breaks. (4) Verify: next scheduled run produces exactly one start line in the job log. | investigation | linux | low | IDEMPOTENT_SAFE | confirm | propose (default) | — | — | — | — |
| fix_ephemeral_container_alerting Guided repair for memory/health alerts on ephemeral `docker compose run` containers (Sowknow guardian HC 2026-07-31: three INCIDENT OPEN pages, INC-20260731-002/005/006, for ghostshell-researcher-run-742d2e58fc28 at 95-100% memory — a one-shot scrape job that no longer existed by the time anyone read the alert; unactionable by construction). Ephemeral run containers match `-run-<hex>` (v2) or `_run_<n>` (v1). Steps: (1) confirm the pattern — `docker ps -a --filter name=run-` and compare alert timestamps to the container's lifetime; if the container is gone, the alert is noise, close it. (2) Fix the MONITOR, not the workload: exclude the ephemeral pattern from alerting while still recording readings (reference: guardian-hc checks/memory.py EPHEMERAL_RUN_RE, 2026-08-02). (3) Do NOT set restart policies or memory limits on run containers to silence the monitor — the container is not the defect. (4) Verify: next patrol records the reading with alert_suppressed, zero pages. | investigation | linux | low | IDEMPOTENT_SAFE | confirm | propose (default) | — | — | — | — |
| fix_stale_monitor_roster Guided repair for monitor-noise drift (Ghostshell watchdog 2026-07-24: 930 alerts on 931 patrols — 'missing containers ghostshell-api ghostshell-v3-dashboard' every 2 minutes — while those containers DO NOT EXIST: the stack was renamed to ghostshell-v1-api / ghostshell-v3-api / ghostshell-dashboard and the watchdog's expected-roster was never updated). A monitor alerting on a stale roster is pure noise and hides real alerts. Steps: (1) get the LIVE roster — `docker ps --format '{{.Names}}'` — never trust repo docs or compose files (two generations coexist; Sowknow's sowknow4- prefix lesson). (2) Diff against the monitor's expected list; update the monitor config to the live names. (3) Verify: next patrol cycle reports 0 missing, alert rate drops to real-events-only. Same fix applies to any healer/probe targeting renamed containers (guardian runbooks, prometheus targets, backup scripts). (4) After ANY stack recreation/rename, re-verify every watcher's roster. | investigation | generic | low | IDEMPOTENT_SAFE | confirm | propose (default) | — | — | — | — |
| handle_llm_timeout_fallback Guided triage for intermittent smoke-test failures on LLM-backed parsing (SOWKNOW search smoke 2026-07-29..08-01: 'intent fallback for French query' — backend log shows 'Intent parsing timed out after 6s ... using fallback'; the fallback defaults language to 'en', which the probe correctly flags. Intermittent, load-correlated, self-clearing). This is DEGRADATION, not outage: the fallback keeps search working with dumber routing. Steps: (1) confirm the signature in app logs — hard parse timeout + 'using fallback' — and correlate failure times with LLM server load (GPU/CPU contention, another tenant's batch job). (2) Do NOT restart the backend: nothing is stuck. (3) Honest fixes, in order: raise the parse timeout past p99 latency, add ONE retry before fallback, cache parsed intents for repeated queries, or pin the LLM to reserved capacity. (4) Keep the smoke test failing on fallback — it is the early-warning for LLM capacity exhaustion; muting it blinds the signal. | investigation | python_app | low | IDEMPOTENT_SAFE | confirm | propose (default) | — | — | — | — |
| investigate_heal_flap Guided investigation for a chronic heal flap (probe:api_feature_health alternating healed/failed/escalated — Ghostshell RCA 2026-07-20 lesson #30). Steps: (1) measure the flap rate in the app's heal-audit table (Ghostshell: SELECT target, result, count(*) FROM v3_heal_audit WHERE created_at > now() - interval '7 days' GROUP BY 1,2); (2) before treating the flap as ACTIVE, check whether the failure counts predate a known fix — a 7-day window spanning the fix date is residual, not active (known instance: the probe:api_feature_health flap was root-caused and fixed 2026-07-21, the GHOSTSHELL_ROLE gate, so counts crossing that date are history — split the window at the fix date; failures after it mean the flap lives on); (3) read the TARGET's own logs around the symptom — adjacent bugs live there; (4) find the app-internal root cause (known: ModelRouter not initialized, so heals never stick); (5) verify feature-level health (Ghostshell: GET /api/v3/health), not just process-alive. Do NOT restart-loop and do NOT duplicate the app's existing healer (I7). The fix is app-side, by the app owner; the platform's job is the measured RCA. Known flap root-cause classes with their own playbooks: a container looping on '[Errno -3] Temporary failure in name resolution' is a DNS/dependency fault -> fix_container_dns_resolution; a monitor/healer alerting on containers that no longer exist is roster drift -> fix_stale_monitor_roster. | investigation | generic | low | IDEMPOTENT_SAFE | confirm | propose (default) | — | — | — | — |
| kill_blocking_query | fix | postgresql | medium | IRREVERSIBLE_RECOVERABLE | confirm | propose (default) | — | — | — | — |
| match_report_rows_to_counts Guided fix for a dashboard/report whose ROWS and COUNTS disagree (SOWKNOW daily report 2026-08-03: incident rows rendered red 'Pending' while the counts above said healed — rows checked `healed`, counts checked `healed or success is True`, and the v2 heal plugin logs `success`). Steps: (1) when the aggregate and its drill-down disagree, the aggregate is usually right and the row rendering is stale — find the condition drift, don't trust either. (2) Give rows the exact same predicate as the counts (reference: guardian-hc daily_report._generate_html, 2026-08-03). (3) Add a unit test asserting row color/status derives from the same function as the counts so the two cannot drift again. Verify: a healed incident renders green 'Healed' in both places. | investigation | generic | low | IDEMPOTENT_SAFE | confirm | propose (default) | — | — | — | — |
| migrate_app_to_least_privilege_role | investigation | postgresql | medium | REVERSIBLE | confirm | propose (default) | — | — | — | — |
| pause_batch_job | fix | generic | low | REVERSIBLE | auto | propose (default) | — | — | — | — |
| plan_host_updates Guided handling of pending system updates (VPS1: 30 pending — Docker, Caddy, apparmor, containerd). Steps: (1) `apt list --upgradable` and split security from routine; apply security updates first. (2) Container runtime updates (docker.io, containerd) RESTART the daemon and every container — schedule a maintenance window and warn the tenant healers will flap during it (expected, self-resolves). (3) Snapshot/backup first when the update touches storage or the runtime. (4) After the window: `docker ps` roster healthy, `needrestart` clean, dashboards green. Do NOT batch-upgrade blind on a production VPS outside a window. | investigation | linux | low | IDEMPOTENT_SAFE | confirm | propose (default) | — | — | — | — |
| probe_timeout_above_slow_threshold Guided fix for a probe whose hard timeout is SHORTER than the slow-path threshold it is meant to detect — making the slow-path branch dead code (SOWKNOW search smoke 2026-08-02/03: 90s HTTP timeout < 120s slow-stream threshold, so a slow-but-working stream was misreported as 'failed: timed out' instead of 'took Ns (>120s)'). Steps: (1) when an alert blames 'timeout' but the service is up, check whether the probe timeout and the slow threshold are even on the same side of each other. (2) The probe timeout must EXCEED the slow threshold it detects — 150s vs 120s (reference: sowknow4 scripts/search_smoke_test.py, 2026-08-03). (3) Add a slow-but-working case to the probe's tests so the branch is exercised. Verify: a deliberately slow stream is reported as 'slow', not 'failed'. | investigation | python_app | low | IDEMPOTENT_SAFE | confirm | propose (default) | — | — | — | — |
| prune_docker_build_cache Guided repair for disk pressure on a docker host that container/image prune cannot relieve (SOWKNOW host 2026-08-03: 72% used, disk_cleanup fired every patrol, image/container prune reclaimed ~nothing — the hog was ~100GB of BuildKit build cache that only /build/prune reclaims; after `docker builder prune -f`: 72% → 50%). Steps: (1) run `docker system df` — if Build Cache dwarfs Images+Volumes, cache is the target. (2) `docker builder prune -f`; BuildKit cache is regenerable — the trade is a slower next build. (3) Automate: a healer that calls /build/prune when disk exceeds warn (reference: guardian-hc healers/disk_healer.py build_cache_prune, 2026-08-03). (4) Rebuild AFTER pruning so you don't rebuild then delete the fresh cache. Verify: df below the warn threshold and build cache no longer the top line of docker system df. | fix | linux | low | IRREVERSIBLE_RECOVERABLE | confirm | propose (default) | — | — | — | — |
| relieve_memory_pressure Guided remediation for sustained memory/swap pressure (VPS1: swap 81% of 8GB with RAM at 48% — the kernel is holding stale pages, or one service leaks slowly). Steps: (1) name the consumers — `smem -rtk` or `for p in /proc/*/status; do ... VmSwap` to rank swap usage per PID, and `docker stats --no-stream` for container RSS. (2) If ONE service holds most swap and is restartable, recycle it in a low-traffic window (a targeted restart, coordinated with the app's own healer — I7). (3) If pressure is chronic across many services, the honest fixes are tuning (vm.swappiness), per-container memory limits, or more RAM — say which; do NOT `swapoff -a` on a loaded host (can trigger the OOM killer). (4) Re-check after 24h: swap refilling fast = a leak to report, not to mask. | investigation | linux | low | IDEMPOTENT_SAFE | confirm | propose (default) | — | — | — | — |
| reload_nginx | fix | nginx | low | REVERSIBLE | auto | propose (default) | — | — | — | — |
| remediate_pg_ssl_exposure | investigation | postgresql | medium | REVERSIBLE | confirm | propose (default) | — | — | — | — |
| restart_celery_worker | fix | python_app | medium | IRREVERSIBLE_RECOVERABLE | confirm | propose (default) | — | — | — | — |
| restart_nginx | fix | nginx | medium | IRREVERSIBLE_RECOVERABLE | confirm | propose (default) | — | — | — | — |
| restart_redis | fix | redis | high | IRREVERSIBLE_DESTRUCTIVE | confirm | propose (default) | — | — | — | — |
| restart_stuck_worker | fix | python_app | medium | IRREVERSIBLE_RECOVERABLE | confirm | propose (default) | — | — | — | — |
| restore_backup_pipeline Guided repair for 'zero backups' drift (Sowknow report 2026-07-24: backup cron exists but target directories are empty — data at risk and SILENT). Steps: (1) prove the failure — run the backup script manually as the cron user and capture stderr (`journalctl -u cron` / `grep CRON /var/log/syslog` for the last runs); a cron that 'exists' but produces nothing is usually failing on credentials, a full disk, or a moved path. (2) Fix the root cause found, then run once by hand and VERIFY a non-empty dump lands (`ls -la <backup_dir>`; for pg_dump, `pg_restore --list` the file — a 0-byte dump is not a backup). (3) Close the silence: backups need a freshness check (alert when newest dump older than 25h) — an unmonitored backup job is indistinguishable from a working one until restore day. (4) Test-restore into a scratch database once; an untested backup is a hope, not a backup. | investigation | linux | low | IDEMPOTENT_SAFE | confirm | propose (default) | — | — | — | — |
| retry_probe_before_false_fail Guided repair for flaky single-shot TCP probes paging false failures (SOWKNOW postgres 2026-08-02: a one-shot connect probe under load paged 'auto-restart disabled — manual action required' while the DB was up and serving; dashboard: PostgreSQL UP). A probe failure is a signal about the probe, not always the target. Steps: (1) before declaring the target down, correlate with other sources — one source UP while the probe says DOWN is a probe problem. (2) Make the probe retry twice with a short backoff (2s) before emitting failure (reference: guardian-hc plugins/infrastructure.py TCP retry, 2026-08-03). (3) Keep real-outage detection: the retry only guards transient TCP backlogs, not connection-refused. Verify: three forced flaps produce zero false pages and a real stop still pages. | investigation | network | low | IDEMPOTENT_SAFE | confirm | propose (default) | — | — | — | — |
| rotate_logs Fix for oversized-text-log drift (e.g. auth.log at 1.5GB — VPS1 health check 2026-07-24). Steps: (1) `logrotate -d /etc/logrotate.conf` to see whether the file is covered — an uncovered file needs a /etc/logrotate.d/ entry (size, rotate count, compress); (2) force one cycle with `logrotate -f <conf>`; (3) for DOCKER container logs (json-file driver) logrotate does nothing — set max-size/max-file in daemon.json or the compose `logging:` block and recreate; (4) confirm disk was reclaimed (`df -h`, `du -sh /var/log`). A log that regrows fast is a symptom (brute-force bursts grow auth.log) — pair with the posture feed. | fix | linux | low | IRREVERSIBLE_RECOVERABLE | auto | propose (default) | — | — | — | — |
| set_container_restart_policy Fix for the docker:no_restart_policy drift class (a container that does not come back after a host reboot, flapping the app's healer into needs_admin — e.g. Ghostshell's v3_heal_audit). Repair: identify the container from the finding evidence, then `docker update --restart unless-stopped <container>`; verify with `docker inspect -f '{{.HostConfig.RestartPolicy.Name}}' <container>`. The container must be on the code-fixed allow-list in executor/command_runner.py — anything else is refused (I6), so an unlisted container is an admin action, not a platform one. EXCLUDE ephemeral one-shot run containers named `*-run-<hash>`: a restart policy is meaningless for a container that runs once and exits, so if v3_heal_audit shows needs_admin storms on docker:no_restart_policy for such names, the defect is in the foreign healer's targeting, not the container — report that, do not 'fix' the container. | fix | linux | medium | REVERSIBLE | confirm | auto | drill-host | 0.0 | 0✓ / 0✗ / 0↩ / 0 flaps / 0 sigs | — |
| set_container_restart_policy Fix for the docker:no_restart_policy drift class (a container that does not come back after a host reboot, flapping the app's healer into needs_admin — e.g. Ghostshell's v3_heal_audit). Repair: identify the container from the finding evidence, then `docker update --restart unless-stopped <container>`; verify with `docker inspect -f '{{.HostConfig.RestartPolicy.Name}}' <container>`. The container must be on the code-fixed allow-list in executor/command_runner.py — anything else is refused (I6), so an unlisted container is an admin action, not a platform one. EXCLUDE ephemeral one-shot run containers named `*-run-<hash>`: a restart policy is meaningless for a container that runs once and exits, so if v3_heal_audit shows needs_admin storms on docker:no_restart_policy for such names, the defect is in the foreign healer's targeting, not the container — report that, do not 'fix' the container. | fix | linux | medium | REVERSIBLE | confirm | confirm | ghostshell-host | 0.0 | 0✓ / 0✗ / 0↩ / 0 flaps / 0 sigs | 2026-07-22 09:12:15 UTC |
| size_heal_verify_to_warmup Guided repair for heal-verify false negatives on warmup-bound services (SOWKNOW 2026-08-02: INC-20260802-001/002 — embed-server restart counted failed because torch model warmup ~60-90s ≫ the 20s+10s verify window; attempts burned, container healthy one patrol later). Steps: (1) when a restart 'fails' verification but the container is healthy a patrol or two later, suspect warmup, not a broken heal — read the service startup log for 'model loaded' timestamps. (2) Size the healer's verify window per service: `verify_delay` ≥ measured warmup and `verify_timeout` ≥ the health endpoint's p99 (reference: guardian-hc core.py:_verify_container_health auto_heal overrides, 2026-08-03 — embed-server 120s/15s, rerank 90s/15s). (3) Re-run the heal on the next flap and confirm it is counted healed on the FIRST attempt. Verify: two consecutive real restarts pass verification without burning attempts. | investigation | generic | low | IDEMPOTENT_SAFE | confirm | propose (default) | — | — | — | — |
| terminate_idle_connection | fix | postgresql | low | IRREVERSIBLE_RECOVERABLE | auto | confirm | ghostshell-db | 0.0 | 0✓ / 0✗ / 0↩ / 0 flaps / 0 sigs | — |
| terminate_idle_connections_bulk | fix | postgresql | medium | IRREVERSIBLE_RECOVERABLE | confirm | propose (default) | — | — | — | — |
| throttle_service_endpoint | fix | generic | medium | REVERSIBLE | confirm | propose (default) | — | — | — | — |
| vacuum_table | fix | postgresql | medium | IDEMPOTENT_SAFE | confirm | propose (default) | — | — | — | — |
Diagnostic trees
| Symptom pattern | Base confidence | Success rate | Coverage | Last updated |
|---|---|---|---|---|
| SSL off while listening on | 0.8 | — | 0 | — |
| active connection as superuser | 0.85 | — | 0 | — |
| api_feature_health | 0.7 | — | 2 | — |
| chronic_foreign_heal: | 0.6 | — | 0 | — |
| config_drift|connections | 0.85 | — | 0 | — |
| dead tuples | 0.8 | — | 1 | — |
| disk_used_pct | 0.7 | — | 0 | — |
| http_5xx_rate | 0.75 | — | 0 | — |
| no_restart_policy | 0.8 | — | 1 | — |
| pg_active_backends | 0.8 | — | 0 | — |
| pg_conn_used_pct | 0.85 | — | 24 | 2026-07-22 10:29:02 UTC |
| pg_idle_in_transaction | 0.85 | — | 0 | — |
| pg_longest_xact_seconds | 0.8 | — | 0 | — |
| pg_replication_lag | 0.7 | — | 0 | — |
| posture_brute: | 0.85 | — | 74 | — |
| process_rss | 0.75 | — | 0 | — |
| queue_depth | 0.75 | — | 0 | — |
| response_time_p95 | 0.85 | — | 0 | — |