Part 4 of a 6-post series on vulnerability management. Part 1 covered why most VMPs fail; Part 2 walked through the scanner architecture; Part 3 was about prioritization. This one is about the SLA layer — the thing that makes the tier real.

Prioritization produces a tier. The tier produces nothing on its own. What turns a tier into actual remediated vulnerabilities is the SLA — a maximum time between ticket creation and resolution, with enforcement that engineering takes seriously.

Most VMPs get the SLA layer wrong in one of two directions. Either there’s no real SLA — tiers are advisory, nothing gets prioritized, the queue just grows. Or the SLA is impossibly strict and gets breached constantly, which is functionally the same as not having one.

The SLA layer I built for a fully custom, budget-constrained VMP I ran end to end sits in the middle: strict durations, documented exception process, and a rollout that got engineering on board instead of around them. Here’s what that actually looks like.


The SLA Shape

Four tiers, four durations:

  • Critical: 3 days
  • High: 7 days
  • Medium: 30 days
  • Low: 90 days

These are the maximum time between ticket creation and resolution. Calendar days, not business days — a vulnerability doesn’t care whether it’s a weekend.

The shape matters more than the exact numbers. A 1-day Critical would be unrealistic for most engineering teams; a 14-day Critical would be too permissive for an actively exploited CVE. The 3/7/30/90 ratio is defensible to leadership (it’s roughly log-scaled, which matches the actual urgency ladder) and tunable in both directions — you can shorten Critical to 48 hours if your organization needs that, or extend Low to 180 days if your backlog is deep.

The thing that trips most teams up is not the number. It’s when the clock starts.


When the Clock Starts (and Why That Matters)

A finding gets created by a scanner. In most VMPs, the SLA clock starts at finding creation — the moment the scanner emits it. That’s a clean rule, and it’s wrong for most organizations that are rolling out a VMP into an environment with pre-existing technical debt.

The problem is that an environment that has never had a VMP has thousands of pre-existing findings. Turn on strict scanning plus strict SLAs on day one, and you’ve instantly put engineering into a position where every single one of their open tickets is already breaching. That’s not a program. That’s a demoralization event.

The design I settled on: the SLA clock starts when a Jira ticket is created, not when the finding is detected. And findings only get promoted to Jira tickets when SecOps flags them for promotion.

The mechanism is a nightly reconciler — an n8n automation that reads findings in DefectDojo, checks the “ready to promote” flags, and creates Jira tickets for the ones that have been flagged. SecOps controls the flag.

This is the single biggest burnout-avoidance decision in the whole architecture. It lets you roll the VMP out gradually against a mountain of technical debt without pretending the mountain doesn’t exist. Day one, maybe fifty findings get promoted. Week two, another fifty. Month three, you’re catching every new finding automatically but still working through the backlog at a pace engineering can actually sustain. The SLA stays strict for every promoted finding, and nothing gets lost — but engineering isn’t drowning in a queue that’s two years old on the day the program turns on.

For environments without heavy pre-existing debt, you can start the clock at finding creation and skip the promotion flag. But if you’re turning on a VMP where there isn’t one already, the gradual promotion model is what makes the first six months survivable.


The Enforcement Loop

An SLA that nobody tracks is a wish. The enforcement loop is what makes it a program.

The automation layer runs in four stages per ticket:

  1. Ticket created, team lead pinged. The nightly reconciler creates the Jira ticket, assigns it based on the owning team, and the team lead gets a notification.
  2. 75% of SLA elapsed — approaching breach. A reminder goes out before the SLA hits.
  3. 100% of SLA elapsed — SLA breached. Breach is logged, team lead re-pinged.
  4. 150% of SLA elapsed — overdue. Escalation signal, visible on the dashboard and in the reporting that goes up the chain.

All four stages post into a dedicated Slack channel that each engineering team has at least one representative in. Those reps are responsible for watching the channel and making sure their team sees new issues. This is important: the Slack reminders are not a replacement for Jira assignment. They’re a redundant notification layer so that an under-watched Jira queue doesn’t silently drop a ticket.

DefectDojo manages the dashboard that shows current open-tickets-vs-SLA status per team, trend over time, and breach rate. Internal automations expose the same data to the security team’s own tracking. Engineering teams can see their own picture; leadership can see the organization-wide picture.

The 4-stage escalation ladder (create → 75% → 100% → 150%) is deliberately gentler than a strict “breach immediately escalates to the VP” model. The intent is to give engineering room to notice and act before the breach lands, and to give them a defined window after breach to still close it before it becomes a reporting event. This is where strict-but-reasonable comes from — the SLA is real, but the system assumes good faith and gives time to respond.


The Exception Process

Not every ticket can be closed within SLA. A vendor hasn’t released a patched version. A fix requires a coordinated migration. The business has decided to accept the risk for a defined period.

The exception process:

  • Anyone can request an exception. The engineer, the lead, the product manager — whoever is closest to the ticket and knows why it can’t close on time.
  • Only SecOps can approve. This keeps the exception bar from drifting. The team requesting the exception doesn’t get to approve its own request.
  • Valid reasons are documented. Primarily: vendor fix pending (there is no patched version to install), and business needs (an explicit decision to defer, signed off by the security head). “We ran out of time” is not a valid reason.
  • Exceptions are logged in DefectDojo. Every exception has an owner, a reason, an extended deadline, and a review date. The dashboard shows active exceptions alongside active tickets so neither disappears.

The two-signature rule for business-need exceptions matters: SecOps approves the exception, but a business-need deferral specifically requires the security head’s sign-off. That’s a deliberate friction. It stops “business need” from becoming a routine escape hatch and keeps exceptions genuinely exceptional.


The Rollout Pattern

The SLAs and the automation would have failed on their own. What made the program land was the rollout.

Before any of this went live, the security team ran a series of meetings with engineering managers across every team that was going to be in scope. Those meetings weren’t an announcement. They were a conversation. The security team walked through the model, the tiers, the proposed SLA durations, the exception process, and — critically — asked for feedback.

Several things changed in those sessions. SLA durations got tuned to actual team throughput. The gradual promotion model (the “clock starts at ticket creation” rule above) came out of a conversation about pre-existing backlog that would have been unmanageable under a stricter interpretation. The Slack rep model came out of a conversation about where engineering was actually paying attention.

Engineering managers who felt consulted became advocates inside their own teams. Engineering managers who felt a program had been dropped on them from above would have become blockers. That’s the whole difference.

Once the program went live, the feedback loop stayed open. The process is still evolving. Metrics get reviewed. SLA durations get re-examined. The exception approval rate gets watched for drift (if SecOps starts approving everything, the exceptions become routine, and the SLA weakens). The channel reps get rotated. None of this is static.

The thing I want to be very specific about: nobody on this program has burned out. Not security, not engineering. The reason isn’t that the SLAs are soft — they’re not. It’s that the whole program was designed to be gradually onboarded, with engineering as a partner in the design, and with an exception process that acknowledges reality instead of pretending it away. Those three things together — gradual promotion, co-designed SLAs, functioning exception process — are what makes a strict SLA layer survivable.


Common Mistakes

Patterns I’ve seen (or avoided by close call) that break SLA programs:

Imposing SLAs top-down. The security team sets the durations, publishes them, and expects engineering to just make it work. This produces either silent non-compliance or active resistance. Co-design the SLAs with the teams who will actually meet them.

Starting the clock on day one for pre-existing findings. If the backlog is deep, every open finding instantly breaches. Engineering gets a dashboard that is 100% red before anyone has done anything. There is no path back from that first impression. Either start the clock at ticket creation with gradual promotion, or exclude pre-existing findings from SLA until they can be triaged.

No exception process. When the only options are “meet the SLA” or “breach the SLA,” engineering will quietly breach constantly. Breaches stop meaning anything. A functioning exception process gives engineering a legitimate way to say “this one needs more time, and here’s why,” which keeps the regular SLA strict.

Exception approval drift. The reverse failure mode: SecOps starts approving every exception that comes in, exceptions become routine, and the SLA becomes advisory. Monitor the exception approval rate. If it’s climbing, the SLA durations are probably wrong — tune the SLA, don’t absorb everything through exceptions.

No Slack / notification layer. Jira assignment alone doesn’t guarantee visibility. A dedicated Slack channel with team reps, with the four-stage notification ladder, catches tickets that would otherwise sit forgotten in a queue nobody is watching.

Metrics only going up the chain. If leadership sees the SLA breach report but engineering teams don’t see their own picture, the program becomes something that happens to engineering, not something they own. Teams need access to their own dashboard — current open tickets, breach rate, exceptions. Transparency inside the team is what makes self-correction possible.


Key Takeaway

An SLA layer is a communication artifact first and an enforcement mechanism second. The durations and the automation matter — Critical 3d / High 7d / Medium 30d / Low 90d, nightly ticket creation, 4-stage escalation, exception process gated by SecOps and the security head.

But the thing that determines whether this program lives or dies is whether engineering was in the room when the SLAs were set, whether the rollout respected the backlog they inherited, and whether the exception process gives them a legitimate way to flag reality.

Set the SLAs strict. Build the automation. Document the exception path. Then do the harder work: get engineering on board before you turn the clock on, and keep listening after you do.

Next in this series: n8n automation patterns — the workflows I use to glue DefectDojo, Jira, and Slack into a VMP that runs without a team of five babysitting it.