23.103. DD 102: systemd service restart policy#
- Design status:
Draft
- Implementation status:
Not started
- DD shepherd:
Florian Dold
- Historical contributors:
Florian Dold
- First published:
2026-09-03
- Last substantive change:
2026-09-03
23.103.1. Summary#
Long-running Taler services should recover from transient failures without operator intervention. They should stop on invalid configuration or another known permanent failure, and repeated crashes must be visible to operators.
The proposed policy retries indefinitely with a fixed ten-second delay. Exit status 6 means that the service is not configured correctly; exit status 9 means that the service encountered another permanent failure. Neither status is restarted.
Systemd 257 in Debian Trixie supports an increasing restart delay, but does
not reset that delay after the service has been healthy for a long time. If a
service reaches a five-minute delay, runs successfully for days, and then
crashes again, the next restart still waits five minutes. Services restarted
regularly through RuntimeMaxSec= also advance the counter. A fixed delay
avoids this persistent state.
23.103.2. Motivation#
Taler’s systemd units currently use different restart policies. Some permit
only five starts in five seconds and some rely on systemd’s default start-rate
limit. Once that limit is reached, Restart=always no longer restarts the
service. A temporary dependency outage can therefore require manual
recovery.
Other units restart after only a few milliseconds or seconds, producing a
tight loop during a longer outage. Some configure RestartSteps= without
the required maximum delay, so the setting has no effect.
Exit handling is inconsistent as well. Native Taler services commonly use status 6 for configuration errors and status 9 for failures that should not be retried, but not all units recognize both statuses and non-C services do not always return them.
An unlimited restart policy also needs monitoring. In particular,
RestartMode=direct skips the failed/inactive transition during automatic
restarts and does not invoke OnFailure= units. A service can therefore
remain in a crash loop without triggering monitoring based only on its current
state.
23.103.3. Requirements#
Long-running services keep retrying after transient failures.
Restarts do not form a tight loop.
Invalid configuration and known permanent failures are not retried.
Exit statuses have the same meaning in every implementation language.
An explicit
systemctl stopstill stops the service.Operators can detect stopped services and crash loops.
The policy works with Debian Trixie.
23.103.4. Proposed Solution#
23.103.4.1. Scope#
The policy applies to long-running product services shipped or deployed by Taler, Anastasis, LibEuFin, Donau, Paivana, Challenger, and related repositories, including distribution-specific copies of their units.
It does not apply to one-shot initialization commands, timer-triggered jobs, garbage collection jobs, or other processes that are expected to finish.
23.103.4.2. Unit policy#
Long-running services use this baseline:
[Unit]
StartLimitIntervalSec=0
[Service]
Restart=always
RestartSec=10s
RestartPreventExitStatus=6 9
StartLimitIntervalSec=0 disables the start-rate limiter. This is needed
because Restart= remains subject to start-rate limiting.
Restart=always restarts the service after clean and unsuccessful exits,
signals, timeouts, and watchdog failures. It does not override an explicit
systemctl stop.
RestartSec=10s avoids a busy loop and has no backoff state to carry across
unrelated failures. Units must not add StartLimitBurst, a non-zero
StartLimitIntervalSec, the legacy StartLimitInterval,
RestartSteps, or RestartMaxDelaySec. A unit may prevent additional
exit statuses when its program documents them as permanent failures.
23.103.4.3. Why the delay is fixed#
Systemd 257 calculates its increasing delay from the number of automatic
restarts. It has no setting equivalent to “reset the delay after ten minutes
of successful operation”. Time spent active does not reduce the counter.
The counter is reset only by an explicit manager action such as
systemctl reset-failed or by stopping and starting the unit.
This matters even for services that rarely fail. For example, after enough failures to reach a five-minute cap:
the service recovers and runs for a week;
it encounters one unrelated transient failure;
systemd still waits five minutes before restarting it.
It is worse for a service using RuntimeMaxSec= as an intentional recycling
mechanism: every scheduled automatic restart advances the same counter, so the
normal restart delay eventually reaches the cap.
Systemd 258 adds RESTART_RESET=1 to its service notification protocol. A
daemon can send this after it considers itself healthy, but systemd 257 ignores
it and every daemon would need explicit support. The policy can be revisited
once the minimum supported systemd version and the applications can provide a
reliable reset.
23.103.4.4. Exit-status contract#
The main process uses these common statuses:
6Configuration is missing or invalid. Retrying the same configuration cannot succeed.
9The service has identified another permanent failure and asks not to be restarted.
All other failures are restartable. Implementations in C, Kotlin, Rust, Go, and other languages must map configuration failures to status 6 rather than a generic status 1. Before returning 6 or 9, the program should log a clear diagnostic.
RestartPreventExitStatus applies only to the main service process. It
does not affect ExecStartPre=. Configuration validation that decides
whether the service should be retried must therefore be reflected in the main
process’s exit status.
23.103.4.5. Application-level retry#
Restarting the process is the fallback when it cannot continue. If a daemon can keep serving useful work or report degraded health while reconnecting to a dependency, it should remain running and retry that operation itself. Domain-specific waits need not use the systemd restart interval.
A service should not wrap its main worker in an internal relaunch loop. Such a loop hides the worker’s exit status and restart count from systemd. If the worker’s termination makes the service unavailable, propagate the result and let systemd restart it.
23.103.4.6. Detecting crash loops#
A service that crashes repeatedly may still appear active whenever monitoring checks it, because systemd keeps restarting it. Checking only whether the unit is active or failed therefore misses crash loops. Deployments must also monitor how often a service restarts. Three restarts within fifteen minutes is a useful default alert threshold.
RestartMode=direct makes failed-state monitoring even less useful: systemd
goes directly from a process failure to a restart without marking the unit as
failed or invoking OnFailure= units. Existing units may keep this setting,
but it is not part of the restart policy.
Services that stop after status 6 or 9 remain in the failed state and should be covered by normal failed-unit monitoring. Journal output is useful for diagnosis but is not an alert.
The relevant behavior is documented in systemd.service(5), systemd.unit(5), and the systemd 258 RESTART_RESET notification.
23.103.5. Test Plan#
Run systemd-analyze verify against the changed units. With a representative
service, check that status 1 keeps restarting at ten-second intervals while
statuses 6 and 9 do not restart. Check that repeated restarts trigger the
deployment’s crash-loop alert.
23.103.6. Definition of Done#
[ ] All checked-in long-running product units follow the baseline policy.
[ ] Distribution and deployment copies match their product units.
- [ ] Service implementations return status 6 for invalid or missing
configuration and reserve status 9 for known permanent failures.
[ ] Worker wrappers do not hide service termination from systemd.
[ ] Supported deployments detect stopped services and repeated restarts.
- [ ] The accepted policy is added to the appropriate developer or operations
reference manual.
23.103.7. Alternatives#
23.103.7.1. Finite start-rate limit#
Systemd’s default limit protects the host and leaves a persistently failing unit in a clear failed state. It also abandons recovery after a temporary outage, so it does not meet the main requirement.
23.103.7.2. Systemd exponential delay#
An increasing delay with a cap reduces load during a long outage. On systemd 257, however, the delay does not reset after a healthy period. This makes an old failure history determine recovery time for a later, unrelated failure. It should be reconsidered when a reliable healthy-runtime reset is available.
23.103.7.3. Timer, wrapper, or external orchestrator#
A timer or wrapper can implement a custom reset rule, but duplicates systemd supervision and complicates stop, status, and exit-code handling. An external orchestrator can provide richer restart policies, but requiring one would change Taler’s deployment model. Deployments that already use an orchestrator may translate the behavior specified here into its native policy.
23.103.8. Drawbacks#
A ten-second delay is not adaptive. A single crash recovers more slowly than
with the shortest current settings, while a long outage causes more attempts
than capped exponential backoff. It also adds ten seconds to an intentional
RuntimeMaxSec restart.
Unlimited retries can consume resources and produce repeated log messages, which is why restart monitoring is required. Incorrectly returning status 6 or 9 can also turn a recoverable failure into one that waits for an operator.
23.103.9. Discussion / Q&A#
23.103.9.1. Why restart after a clean exit?#
These units represent services expected to remain available. Processes that are expected to finish belong in one-shot or timer-driven units.
23.103.9.2. Why not treat every startup failure as permanent?#
Startup may depend on a database, network service, or credential agent that is temporarily unavailable. Only failures classified by the program as invalid configuration or permanent should suppress restart.