Technical Guide · Linux and System Administration
Systemd Dependencies: Design with After, Wants and Requires
After controls ordering and does not automatically start the other unit. Wants is a weaker requirement, Requires a stronger one; ordering is still separate. Starting A after B does not…
Technical review:
Architecture and operating model
After controls ordering and does not automatically start the other unit. Wants is a weaker requirement, Requires a stronger one; ordering is still separate. Starting A after B does not prove B is application-ready.
Network-online.target can wait for manager-defined network readiness during boot; it does not continuously guarantee remote-service reachability. Applications need bounded retries/readiness when DNS or databases are unavailable. Blind sleeps hide dependency defects.
Oneshot, simple and notify differ in readiness semantics. Use drop-ins and inspect effective units after daemon-reload. Resolve cycles by redesigning the minimum relationship rather than adding more Requires directives.
- 1Unit requirement
- 2Ordering graph
- 3Process/readiness
- 4Retry/health
Design parameters
- Activation need
- Choose activation separately from ordering; select Wants/Requires by the actual requirement.
- Readiness
- Process existence, listening ports and application health are different.
- Retry limit
- Coordinate RestartSec and application retry budgets without creating restart storms.
Worked example
An API that should follow PostgreSQL but can return controlled 503s during outages may need ordering/retry rather than a mandatory stop relationship. Delay database startup by 20 seconds in a lab and measure process start, readiness and first successful query separately.
Example commands: replace lab values and confirm permissions and software versions before use.
systemctl cat lab-api.service
systemctl show lab-api.service -p After -p Wants -p Requires
systemd-analyze critical-chain
journalctl -u lab-api.service -b
Troubleshooting
| Observation | Likely cause / distinction | Verification |
|---|---|---|
| After set but DB not started | Ordering is not activation. | Inspect Wants/Requires and activation transactions. |
| Cycle detected | Circular ordering relationship. | Inspect units together with drop-ins. |
Acceptance checks
- Inspect the effective unit.
- Separate ordering/activation.
- Test application readiness.
- Measure boot delays.
- Test outage retry behavior.
- Verify drop-in rollback.
Related concepts
Dependencies and restart order
Services commonly depend on identity, DNS, time, networking, databases and licensing. Record a dependency graph describing conditions for operation, not merely an equipment list. Recovery order follows that graph; circular dependencies may require emergency access paths. Distinguish restored infrastructure from resumed business activity. Assign an owner, validation method and alternative access path to each dependency. Test assumptions by deliberately making one component unavailable in a controlled end-to-end exercise.
Service lifecycle and dependencies
A running process does not prove that users receive correct service. Review startup order, network or database dependencies, environment variables, file access and health checks together. Restart loops can conceal the root cause. Compare exit codes, recent logs and resource limits, and identify whether systemd or container management triggered the restart. After a controlled restart, validate sessions, writes and dependent services.
Availability versus recovery
High availability aims to keep service running through specified failures with a short interruption; backup recovers lost or corrupted data from an earlier point. A cluster can replicate an accidental deletion to another node. HA therefore does not replace backup. Consider DNS, identity, network, storage and power dependencies together. Successful node failover is insufficient by itself: measure user sessions, application writes and external integrations after the transition as well.
Version and support lifecycle
Installability does not prove production support. Review the compatibility chain across OS, application, drivers, extensions and management tools. Update plans should record version, support end, restart needs and rollback methods. An unrepresentative test environment can produce misleading results. Validate service health and existing workflows after a change, not just version numbers. Remember that pinning a version can also prevent future security fixes.
Telemetry and time correlation
Telemetry combines logs, metrics and events that explain system behaviour. A log describes an event, a metric shows behaviour over time, and a distributed trace follows a request across components. Clock differences can make one event appear to occur at several times. Use synchronized clocks, reliable source identifiers and consistent time-zone handling. Alarm design should consider duration and user impact alongside thresholds. Monitor gaps in collection separately: absence of logs must not be interpreted as absence of incidents.
File permissions and service identities
File access combines users/groups, permission bits, ACLs and security policies. Running an application as root can hide permission problems; prefer a justified, restricted service identity in production. Directory traversal permission differs from file-read permission. Sharing permissions do not necessarily override filesystem restrictions. Identify the actual runtime user, directory chain and ACLs when troubleshooting. Test narrowly required access rather than opening permissions globally.
Primary documentation
Prepared by the Doz Teknoloji technical team using the primary references below. Calculations and lab scenarios state their assumptions; validate the applicable product version before rollout.