Technical Guide · Cybersecurity
Service-to-Service mTLS: Identity, Trust Chains and Authorization
mTLS authenticates both server and client certificates during a TLS connection. Service A may authenticate to a payment service without being authorized to perform every payment…
Technical review:
Architecture and operating model
mTLS authenticates both server and client certificates during a TLS connection. Service A may authenticate to a payment service without being authorized to perform every payment operation. A separate authorization decision maps certificate identity to application permissions.
When TLS terminates at a load balancer, the backend may not see the original client certificate. Forward verified identity over a trusted channel and remove identically named external HTTP headers. Otherwise forged identity headers may cross the trust boundary.
OAuth client authentication and certificate-bound access tokens are distinct mechanisms. The latter requires the resource server to validate the token certificate binding too. During renewal, evaluate old/new identities, existing connections and token lifetimes together.
- 1Client certificate
- 2TLS chain validation
- 3Service identity and role
- 4Application decision
Design parameters
- SAN identity
- Match the service name or URI identity explicitly; a valid certificate chain alone does not grant application access.
- Trust root
- Restrict CA lists to required issuers for client and server; avoid automatically trusting every enterprise root.
- Key protection
- Keep private keys out of images. Assign ownership for workload identity, file permissions and rotation.
- Revocation and renewal
- Short lifetimes, revocation checks and emergency termination are separate needs; test CRL/OCSP behavior in the deployed client version.
Worked example
Use orders.example.test and an inventory-client identity in a lab. Permit only GET /stock for the trusted client; POST /payment must be denied by application policy even after TLS succeeds.
During a one-hour renewal pilot, observe the connection pool. A new certificate may work for new connections while old connections remain active. Record revocation results for both existing and new sessions.
Example commands: replace lab values and confirm permissions and software versions before use.
openssl s_client -connect orders.example.test:443 -servername orders.example.test -CAfile ca.pem -cert client.pem -key client.key -verify_return_error
Troubleshooting
| Observation | Likely cause / distinction | Verification |
|---|---|---|
| unknown ca | Missing intermediate or incorrect trust store. | Compare the presented chain with the actual client CA file. |
| TLS succeeds, HTTP 403 | Identity authenticated but role not mapped. | Correlate SAN identity and policy decision using one request ID. |
| Outage after renewal | Stale root, cache or connection pool. | Repeat the request with a fresh process and connection. |
Acceptance checks
- Test authentication and application authorization separately.
- Reject a client without a certificate.
- Exercise an untrusted CA.
- Reject mismatched SAN identity.
- Verify a pilot request during key renewal.
- Record identity and decision reason in audit events.
Related concepts
TLS and certificate validation
TLS protects confidentiality and integrity in transit; certificate validation helps verify the peer’s identity. Evaluate names, chains, validity periods and trusted roots together. Encryption does not prove correct application authorization. If a reverse proxy or inspection device is used, show where TLS terminates. Disabling validation is not a permanent troubleshooting solution: investigate hostname mismatch, missing intermediate certificates and incorrect device clocks separately.
Authentication and sessions
Authentication proves who a user or workload is; authorization determines what that identity may do. Successful sign-in does not grant access to every resource. User sessions, service identities, API tokens and device certificates have different lifecycles. Design session duration, token renewal, employee departure, lost-device handling and emergency access alongside initial sign-in. Measure which existing sessions remain usable and which new accesses are denied when the identity provider becomes unavailable.
Trust boundary and failure domain
A trust boundary separates components governed by different access decisions; a failure domain groups resources that one event can affect together. Two VLANs do not create a strong trust boundary when routing between them is unrestricted. Backups in different folders still share a failure domain if one administrator can delete both. Assess physical location, identity provider, management account, network path and power source separately. Test which access paths and recovery options remain available when a component is lost.
Least privilege and separation of duties
Least privilege limits the permissions needed for a task to specific resources and time periods. Giving everyone the same administrator role complicates access reviews and investigations. Separate daily-use identities from privileged accounts, and restrict interactive use and unnecessary network access for service identities. An access matrix should show which operation an identity may perform on each resource. Test whether an application keeps working after a permission is removed: unnecessary privileges sometimes become visible only during controlled reduction.
Version and support lifecycle
Installability does not prove production support. Review the compatibility chain across OS, application, drivers, extensions and management tools. Update plans should record version, support end, restart needs and rollback methods. An unrepresentative test environment can produce misleading results. Validate service health and existing workflows after a change, not just version numbers. Remember that pinning a version can also prevent future security fixes.
Telemetry and time correlation
Telemetry combines logs, metrics and events that explain system behaviour. A log describes an event, a metric shows behaviour over time, and a distributed trace follows a request across components. Clock differences can make one event appear to occur at several times. Use synchronized clocks, reliable source identifiers and consistent time-zone handling. Alarm design should consider duration and user impact alongside thresholds. Monitor gaps in collection separately: absence of logs must not be interpreted as absence of incidents.
Primary documentation
Prepared by the Doz Teknoloji technical team using the primary references below. Calculations and lab scenarios state their assumptions; validate the applicable product version before rollout.