Article

What service levels make sense for business-critical automations?

If an automation has become part of how the business operates, “it usually works” is not a useful reliability standard. The right service level starts with the business outcome, not a generic uptime percentage.

By Qwaname Kenobisan·Business Automation·2 September 2026

Automations often begin as convenience tools. A workflow copies information, sends a notification or moves a request to the next stage. Over time, more work depends on it. The business stops performing the old manual check because the automation is expected to run.

At that point, reliability becomes an operating requirement. The problem is that many organisations still manage the workflow as though it were a small script. There may be no agreed definition of success, no response expectation, no recovery target and no owner watching whether failures are accumulating.

A service level is useful only when it describes the reliability the business outcome actually needs.

Start with the consequence of failure

Not every automation deserves the same level of engineering or support. A workflow that posts a weekly internal summary can tolerate more delay than one that routes paid customer orders, updates financial records or provisions access for a new employee.

Classify the automation by consequence. Ask what happens if it fails for five minutes, one hour, one business day or several days. Consider customer impact, financial impact, compliance exposure, operational backlog and the effort required to reconstruct missed work.

This creates a rational basis for investment. Reliability should be proportional to business criticality rather than the technical sophistication of the workflow.

Define the service in terms of outcomes

Traditional infrastructure often measures availability: was the service reachable? Business automation needs a wider definition. A workflow can technically run while still producing the wrong result, skipping records or leaving a downstream system incomplete.

A useful service definition may include several measures:

  • Execution success: the percentage of eligible workflow instances that complete correctly.
  • Timeliness: how quickly work must begin or finish after the trigger occurs.
  • Correctness: whether the intended business state was actually reached.
  • Recovery: how quickly failed or delayed work must be restored.
  • Backlog visibility: whether stuck work can be identified before it becomes customer impact.

The point is not to create a complicated scorecard. It is to measure what the business would notice if the automation stopped doing its job.

Use targets, not promises you cannot support

Google's Site Reliability Engineering guidance distinguishes service-level objectives from vague reliability expectations and uses error budgets to make trade-offs explicit. The same principle can be adapted to business automation without pretending every workflow is an internet-scale service.

For a critical workflow, the organisation might define an objective such as: 99.5% of eligible transactions complete correctly within fifteen minutes over a rolling month, with priority failures acknowledged within thirty minutes during supported hours. Another workflow may only need same-business-day completion.

Do not choose 99.9% because it sounds professional. Higher targets require better monitoring, redundancy, support coverage, testing and recovery mechanisms. Reliability has a cost.

Separate detection from recovery

An automation can fail for hours if nobody knows it failed. This is why service levels must include detection. A useful operating model defines how quickly a meaningful failure should become visible and how quickly somebody should begin responding.

The article Who Watches the Automations? makes the broader observability case. Service levels add a sharper question: once the failure is visible, what response does the business actually require?

For important automations, define at least four time expectations: detection, acknowledgement, containment and restoration. These do not need to be contractual SLAs. They can be internal operating objectives that guide support priorities.

Design for retries without creating duplicate work

Recovery is not simply “run it again.” A retry can duplicate invoices, notifications, tasks or downstream records if the workflow was only partially successful. Service-level design therefore depends on technical controls such as idempotency, correlation IDs, checkpointing and reconciliation.

The reliability target should be supported by a recovery path that is safe. If failed work has to be reconstructed manually, document the procedure and ensure the necessary audit data is retained.

Give criticality levels different support models

A practical approach is to create a small number of tiers rather than negotiate a unique standard for every automation.

  • Tier 1 — business critical: direct customer, revenue, financial or control impact; active monitoring; rapid escalation; tested recovery; named primary and backup owners.
  • Tier 2 — operationally important: material internal delay if unavailable; monitored during business hours; documented recovery and ownership.
  • Tier 3 — convenience: limited consequence; failure can be handled through routine support without immediate intervention.

The exact names do not matter. The discipline does. Criticality should influence how much engineering, monitoring and support the workflow receives.

Measure the service level from the user's side

Automation platforms may report that a run succeeded, but the customer or employee may still be waiting because a downstream step did not occur. Measure the workflow as close as possible to the business outcome.

For example, an order-processing automation should not be judged only by whether the orchestration flow returned a success status. Check whether the order was created correctly, whether downstream fulfilment received it and whether exceptions were surfaced.

This is the difference between platform health and operational health.

Review targets as the business changes

A workflow that begins as a low-volume internal convenience can become critical as usage grows. Conversely, an old automation may become less important after a process changes. Review service levels when transaction volume, customer dependency, operating hours or architecture changes.

The target should describe the current business dependency, not the importance the workflow had when it was first built.

What better looks like

A mature automation capability knows which workflows are critical, what success means for each one, how quickly failures must be detected and recovered, and who owns the response. Reliability expectations are explicit enough to guide engineering and support decisions without pretending every workflow needs maximum uptime.

If the business cannot answer what happens when a critical automation fails tonight, the missing control is not another dashboard. It is an operating agreement around the automation.

Related Mellorca servicesManaged Digital Operations can help establish monitoring, support, change and recovery disciplines for important workflows. Business Automation & AI Workflows can design reliability into new automations from the start.

Sources and further reading