By the end of this you will know which of four events deleted your self-hosted runner's registration, and you will have a one-line test that says so on every host in the fleet. The line that sends people here is followed by an exit code of 0, so the machine looks healthy and nothing brings the runner back.

Prerequisites

  • Shell access on the runner host, or kubectl logs on the runner pod for ARC.
  • The install directory, ~/actions-runner by default, with its _diag folder intact.
  • gh authenticated with admin:org (org runners) or repo admin scope.
  • Runner v2.337.0, published 2026-08-26 and still the newest release as I write this on 2026-09-28. The source quoted below is from that tag.

If your symptom is a runner that accepts a job and then sits on it, you want the stuck-slot walkthrough for "Acknowledging runner request" instead. This page is for the runner that goes away.

Step-by-step

1. Read which sentence you actually got

There are two near-identical strings in Runner.Listener, from two source files, and they have opposite consequences for your disk. Neither is documented on docs.github.com.

grep -h "no longer exists" ~/actions-runner/_diag/Runner_*.log | tail -5

MessageListener.cs, inside IsSessionCreationExceptionRetriable(), prints on a TaskAgentNotFoundException during session creation, before the runner ever reaches Listening for Jobs:

The runner no longer exists on the server. Please reconfigure the runner.

The method returns false, the runner stops, and your config files survive.

Runner.cs, inside RunAsync, prints on the same exception class after the listener is up and polling (source at v2.337.0):

catch (Exception ex) when (ex is TaskAgentNotFoundException || ex is RunnerNotFoundException)
{
    Trace.Info($"Runner registration no longer exists while retrieving messages. {ex.Message}");
    _term.WriteError("The runner no longer exists on the server. Cleaning up local configuration.");
    skipSessionDeletion = true;
    cleanupLocalConfigAfter404 = true;
    break;
}

cleanupLocalConfigAfter404 is consumed further down the same method:

if ((settings.Ephemeral && runOnceJobCompleted) || cleanupLocalConfigAfter404)
{
    configManager.DeleteLocalRunnerConfig();
}

DeleteLocalRunnerConfig() in ConfigurationManager.cs calls _store.DeleteCredential(), keyManager.DeleteKey() and _store.DeleteSettings(), printing Removed .credentials and Removed .runner. Then RunAsync returns Constants.Runner.ReturnCode.Success.

The Trace.Info string Runner registration no longer exists while retrieving messages. appears only on the Runner.cs path. Its presence in _diag proves the config was deleted, long after the terminal scrollback is gone.

2. Check what is left on disk

ls -la ~/actions-runner/.runner ~/actions-runner/.credentials 2>&1

Both files missing means the 404 arrived mid-listen and the runner cleaned itself out. Both present with a dead service means somebody or something stopped the unit, which is a different problem with the same alert.

3. Ask the API what the registration looked like

gh api /repos/OWNER/REPO/actions/runners \
  --jq '.runners[] | [.id, .name, .status, .busy, .ephemeral, .version] | @tsv'

The ephemeral field separates a reaped JIT worker from a deleted long-lived one. The version field arrived with the 3 September 2026 Actions update and is documented in the self-hosted runner REST reference; if your GHES version does not return it, fall back to reading .runner on each host.

4. Look for a human or an API call

gh api '/orgs/ORG/audit-log?phrase=action:org.remove_self_hosted_runner&include=all' \
  --jq '.[] | [.created_at, .actor, .action] | @tsv'

Check the action names against your own audit-log event list before you trust an empty result, since the event set differs between Enterprise Cloud and Server. Drop the phrase= filter and grep the output for runner if nothing comes back. The automatic offline sweep does not reliably surface as an audit event, so absence of a row is not evidence that nobody deleted anything.

5. Match your evidence to one of four causes

CauseWhat you seeSeparating checkFix
JIT/ephemeral runner reaped while idleListening for Jobs, then the line 2 to 6 minutes later, no job ever dispatchedephemeral was true in the REST listing; the gap between the two log timestamps is minutesMint the JIT config as late as possible and treat the exit as "replace the worker", not "retry". Track issue #4726
Deleted by a human or an API callArbitrary uptime, no patternAudit log shows a removal event for that runner nameFind the caller and stop it deleting live runners
Same name re-registered elsewhere with --replaceTwo hosts, one name; the older process diesA fresh registration event for a name you already hadMake runner names unique per host
Offline sweepHost was down for days before the restartGitHub's removal policy: 14 days for normal runners, 1 day for ephemeralRe-register, the token is gone
Runner log goes quietWhich sentence?404 at session creationMessageListener.cs.runner survives404 mid-listenRunner.cs RunAsync.runner and .credentials deleted, exit 0Registration intactcheck runner versionWas it ephemeral or JIT?Reaped before a job arrivedissue 4726Deleted, replaced, or swept offlinecheck the audit log"Please reconfigure the runner""Cleaning up local configuration""no line at all, still Listening for Jobs"yesno

6. Understand why systemd sat on its hands

The unit template shipped at actions-runner/bin/actions.runner.service.template carries no restart policy:

[Service]
ExecStart={{RunnerRoot}}/runsvc.sh
User={{User}}
WorkingDirectory={{RunnerRoot}}
KillMode=process
KillSignal=SIGTERM
TimeoutStopSec=5min

With no Restart=, systemd defaults to Restart=no. Pair that with an exit code of 0 and the unit lands in inactive (dead) with status=0/SUCCESS, which most alerting cannot tell apart from an operator running svc.sh stop. Issue #3909, opened 2025-06-14 and closed as not planned, is where GitHub declined to change that default.

7. Make the exit loud, per runner type

For a long-lived runner, add the drop-in and pair it with re-registration, because on the Runner.cs path the credentials are already gone and a bare restart loop has nothing to authenticate with:

# /etc/systemd/system/actions.runner.<name>.service.d/override.conf
[Service]
Restart=always
RestartSec=15

For JIT, shorten the idle window instead of restarting anything:

gh api -X POST /repos/OWNER/REPO/actions/runners/generate-jitconfig \
  -f name="runner-$(uuidgen)" -F runner_group_id=1 \
  -f 'labels[]=self-hosted' -f 'labels[]=linux'

Verify it works

Run this on every runner host. It costs nothing and it is offline:

test -f ~/actions-runner/.runner || echo "CONFIG DELETED: 404 path, see _diag"

Then confirm the service state matches the story:

systemctl show 'actions.runner.*' -p Id -p ActiveState -p ExecMainStatus -p ExecMainExitTimestamp

ActiveState=inactive with ExecMainStatus=0 and a missing .runner is the fingerprint of the mid-listen 404. An operator stop leaves .runner in place, so the file separates the two cases with no guessing. After the drop-in, kill the runner process and expect ActiveState=active again within RestartSec=15.

Common pitfalls

Assuming the removal policy explains a six-minute death. The documented sweep is 1 day for ephemeral runners and 14 days for normal ones. A JIT runner that vanished minutes after registering is not covered by that policy, and the behaviour is still open as issue #4726, filed 16 September 2026 with this exact log. Quoting the 14-day rule at that symptom sends people rebuilding hosts for no reason.

Reaching for Restart=always on its own. On the cleanup path the config is deleted before the process exits, so the restarted runner immediately fails to authenticate and burns a restart budget every 15 seconds.

Confusing the two sentences. "Please reconfigure the runner" leaves your files alone; "Cleaning up local configuration" does not. Grep for the exact tail of the line, not for "no longer exists".

Blaming the 25 September minimum-version enforcement. Full enforcement for GitHub Enterprise Cloud began 25 September 2026 (30 July 2026 for Data Residency): registration requires 2.329.0 or later, and each release must be installed within 30 days of publication or the service stops queuing jobs to that runner (changelog). Do the arithmetic: 2.337.0 published 2026-08-26, so its 30-day window closed on 2026-09-25, the same day enforcement went full. Anyone still on 2.336.0 today is outside it. Enforcement prints nothing. It leaves a runner online in the UI, parked at Listening for Jobs, receiving no work, in the same way a kubelet flag removal leaves nodes that never join with a perfectly healthy-looking process. Audit the fleet:

gh api /orgs/ORG/actions/runners --paginate --jq '.runners[].version' | sort -u
gh api /orgs/ORG/actions/runners/deprecations/2.336.0

The deprecations lookup returns runner_version, runtime_deprecates_at and registration_deprecates_at. It is version-scoped and will not enumerate your fleet, which is why the first command exists. The repository-level form is /repos/{owner}/{repo}/actions/runners/deprecations/{version}.

Treating this as the only September surprise. The node20 shim now running on Node 24 hit the same fleets in the same window, and neither change required anyone to touch their workflows. If you are auditing runner supply chain at the same time, the gaps a SHA pin leaves open are worth a pass while you are in there.

Wrap-up

Two facts carry this whole diagnosis: the cleanup branch fires only after the listener is up, and it returns success. Put the .runner test in your config management so a deleted registration raises an alert on the next run, then sweep versions before the next release starts its 30-day clock:

gh api /orgs/ORG/actions/runners --paginate --jq '.runners[].version' | sort -u