Practical cybersecurity from Northern Ireland

MFA Rollout Planning: The Hidden 80% Nobody Warns You About

Recently I attended an event in Belfast, and the room was filled with folk with far more IT experience than I have. Most were IT managers or team leaders. Very few were security specialists or, more importantly, had any security specialist on their teams. Almost all of them were responsible for it, though, and were discussing the recent news that Microsoft is retiring its own SMS and voice authentication services.

Chatting to them, it became apparent that to most this was a bit of a surprise and had snuck up on them. It was released during the holiday period (the July week here in Northern Ireland is when most are off), and they found out either through being assigned the task via the Microsoft Message Centre, or in some cases having a senior leader read an article on social media or the news and get in touch to find out “what are we doing about this?”

The thing I noticed was that all their questions were technical. Which authentication method to choose. How Conditional Access policies work together. What might break when legacy authentication is removed.

As it was a predominantly public sector group, I realised this would be their first time, or at least in a long time, of rolling out an MFA change. With most still primarily using SMS, this was going to be a significant culture change that requires more than an email update and clicking some buttons in Entra. That’s the easy part.

They had skimmed the guides, read the social media posts and understood the material, but their concerns were focused on the wrong 20%…..which was evident when I started asking about change control and service accounts.

This isn’t a criticism of them or the guides. The articles I’ve read about the retirement are solid and explain how to configure it well. But they’re written by technical people who are used to writing for other technical people, usually starting and ending at the configuration step. That leaves the bigger challenge unexplored.

So I sat down to write an article on how to roll out MFA, and it ended up being a monster several thousand words long, which I’ve now broken into a series. In it, we’ll look at the real factors that determine a rollout’s success, which are different from what most people expect: understanding what you’re changing, agreeing on business needs, preparing the people involved, and planning for problems.

That’s where most of the work happens, yet it’s rarely written about.

Setting up MFA is just the easy 20%

Like most people I’ve spoken to, the first time I rolled out a project like this I thought I’d spend most of my time in Entra on Conditional Access policies and authentication methods. Instead, most of it went on conversations and documentation.

Conversations about people who were on leave just before go-live, staff working overseas worried about connectivity, and executives who didn’t want any disruption. Conversations with colleagues about potential issues, planning so as to avoid impacting others’ work, and explaining the why as well as the how. Then producing the documentation, guides that meet the needs and knowledge level of the people it’s being rolled out to, project planning docs and DPIAs. All of that before you even look at user accounts being used as service accounts with no clear owner, or what they actually do.

You won’t find any of that in Microsoft’s documentation, but these issues determine whether the rollout succeeds or becomes a total nightmare. MFA projects rarely fail because someone can’t find the Conditional Access settings. They fail when something that should have been discovered, agreed, communicated or tested beforehand gets missed. The outcome is usually decided weeks before the first authentication prompt, not on the day the policy goes live.

TIP: Book your change advisory board slot before you start building anything. In public sector the CAB lead time is often weeks, and it’s usually the dependency people discover last.

Nobody should do this in one go

Before anything else, let’s be clear on this: there is no single switch-on date, and there should never be one.

Every rollout I’ve run has gone out in waves. A small pilot first, then progressively larger groups, usually split by department or by risk, with a deliberate gap between each one to fix whatever the last wave taught you. The gap is the point. A wave you don’t learn from is just a slower big bang. I’m sure we’ve all either done it or helped a colleague recover from a change that brought down production….. You know who you are!

Phasing does three things that a single cutover cannot. It caps your blast radius, so a mistake costs you 30 confused users instead of 3,000. It gives your service desk a survivable support curve rather than one catastrophic Monday. And it means every group after the first benefits from problems the earlier groups found for you.

It also changes what rollback means. Pulling back a wave of forty people is an inconvenience. Pulling back the entire company is an incident, and by then you’re making the decision in front of an audience.

TIP: Never enforce a wave on a Friday or the day before a bank holiday. Your support curve peaks 24 to 48 hours after enforcement, and you want a full working week to absorb it. Ask each department for their blackout dates, month-end for finance, enrolment for education, before you build the schedule rather than after.

This matters for how you read the rest of this article. The six steps below aren’t a straight line you walk once:

  • Discover and Define happen up front, once, across the whole estate.
  • Prepare, Pilot, Protect and Prove repeat for every wave.

Your first wave is genuinely a pilot. Your second is a pilot with a bigger sample. You’ll stop calling them pilots at some point, but you should never stop treating them like one.

A six-step approach

I now think of MFA readiness in six steps: Discover, Define, Prepare, Pilot, Protect and Prove. I’ve refined this over time by making the mistakes that informed my experience, and the PTSD that came with them.

None of these steps is about choosing an authentication method. They’re about making sure the method you choose actually works for real users, who will complain loudly to their managers/everyone in earshot if you get it wrong.

Discover what you’re actually changing

Before changing any policies, find out what’s actually in your environment, not just what you think is there. Look for stale accounts, guest accounts, old user accounts being used as service accounts, unexpected legacy authentication, and applications that might break under the new policies.

Don’t let an enforcement date be the day you discover these issues.

This matters because if you find a user account being used as a service account, an SMTP AUTH dependency, or an app authenticating with a user’s credentials on enforcement day, you’ll be troubleshooting live instead of preparing the next wave. Run reports on legacy authentication sign-ins and unregistered users weeks before rollout. Microsoft’s sign-in logs and Conditional Access “what if” tool will surface most of this early. I also used Maester for some of the auditing work, which is a genuinely useful tool.

Do this once for the whole estate up front, then re-check per wave. Every department has its own strange application that nobody mentioned in the kickoff.

TIP: Entra sign-in logs only retain 7 days on the free tier and 30 days on P1/P2. If you want a proper baseline, export to Log Analytics before you start, or you’ll be making decisions on a partial picture. Report-only mode on a Conditional Access policy also gives you genuine impact data without enforcing anything, so run it for at least a full month to catch the monthly processes.

Define the requirements in four parts, not one

Saying “roll out MFA to everyone” is a headline, not a requirement. In my experience, it breaks down into four separate areas, but most teams only document the first.

Security requirements: what must be in place. Block legacy authentication, use phishing-resistant methods for privileged accounts, and plan to retire SMS eventually, even if not right away.

Business requirements: minimise disruption, support overseas and remote staff, and allow for exceptions in genuine business situations instead of pretending they don’t exist.

Support requirements: make sure the service desk is prepared, with user guides, troubleshooting steps and a known escalation process.

Governance requirements: an approved exceptions process, clear ownership, and documentation you can produce during an audit.

At the core, the questions are simple. Who needs access, to what, from where, under what conditions, and what happens when their usual method isn’t available? You don’t need a certification to ask these, but you do need to ask them before building a policy, not after someone is locked out.

Getting approval from senior leadership is generally easier than most teams expect, once you stop focusing on technical detail. Leaders care about risk and accountability and want solutions, not lectures. Say “we need to reduce the risk of account compromise and the disruption it causes” rather than “we need phishing-resistant authentication methods”, and sign-off gets a lot easier.

TIP: Get the risk owner named in writing at sign-off. “The business accepted the risk” is worthless in an audit without a name and a date attached to it.

Prepare the people, not just the policies

The security control might be the same for everyone, but the rollout can’t be.

A technical team, a general office worker, a frequently travelling executive/salesperson, and someone who rarely uses IT might all need different documentation, timing, and support. Assuming “it’s easy, they’ll work it out” is where a lot of rollouts start losing time they never get back.

This is also why your wave order matters. Sequencing by department or risk rather than alphabetically lets you match the level of preparation to the group you’re about to hit, and put the groups who need the most hand-holding later, once your guides have been through a few rounds of real use.

The main difference between a smooth wave and a difficult one usually comes down to three things: an accessible guide covering nearly every possible outcome, having the Microsoft Authenticator installed before their go-live date and getting the message out that everyone needs to read the guides (hardest one of all). Most real problems aren’t technical. They come from lack of preparation, which is a communications problem rather than a configuration one.

Communication isn’t a single email. It’s an ongoing effort using intranet posts, department-specific go-live dates, reminders and manager briefings. Tsedal Neeley and Paul Leonardi’s research on this is blunt about it: to get people to actually do something, managers need to ask them at least twice, and the managers who got things done fastest deliberately repeated themselves across more than one channel. Good communication answers five questions clearly: why this is happening, what needs to be done, by when, what happens if nothing is done, and what to do if it isn’t working. This is where knowing your marketing and comms teams is invaluable. They are the experts here; they do this all day, every day and know the audience better than anyone, so use them and listen to them.

Directors and senior managers play a bigger role than most security teams realise. People listen to IT, but they pay more attention to the big bosses. If department heads aren’t informed and supportive, the rollout gets much harder before any policy changes in the background.

TIP: Don’t put a working QR code in your guide. People scan the one on the paper instead of the one on their screen. Use a greyed-out placeholder image with a caption saying “your screen will show a code here”.

Pilot to test your assumptions, not your MFA

A pilot isn’t there just to prove MFA works technically. It’s there to find out what you missed while it’s still cheap to find out.

Choosing the IT department as your pilot group tells you very little, since they’re the least representative group you have. A good pilot spans different technical skill levels, device types, locations, remote workers, and at least one executive. I almost always identify the whingers and moaners in advance and include them deliberately, because they give the most direct feedback, no matter how painful their delivery.

TIP: Include one person about to go on leave and one who has just come back. Both surface edge cases nobody plans for: re-registration after time away, and Temporary Access Passes that expired while they were off.

The pilot isn’t really about testing MFA. It’s about testing your assumptions. Instructions you thought were clear get misunderstood. Steps you assumed people would follow get skipped. Documentation that seemed complete has gaps you only notice when someone outside the project team tries to use it. You’d think it would be obvious not to scan the QR code printed in the guide when there’s one on the screen in front of you….. but it happens more often than I can say.

Every wave after the pilot should have the same feedback loop attached, even when the numbers get bigger. The moment you stop collecting feedback is the moment you go back to running a big bang in instalments.

I use Microsoft Forms for this, asking 8-10 questions to get as much feedback as possible for the next rollout phase.

Sample questions for feedback form

  1. How would you rate your overall experience with the new sign-in process? (radio button)
  2. How easy was it to set up on your device? (radio button)
  3. How easy was it to use the new sign-in method as part of your normal working day? (radio button)
  4. How clear and helpful were the instructions and communications provided? (radio button)
  5. Did you experience any issues or challenges during the setup or use? (Yes/No radio button, with a follow-up text box that appears if they answer Yes)
  6. Has the new sign-in process had any impact on your ability to carry out your work? (radio button)
  7. If you experienced an impact, can you please describe it (text box)
  8. What’s one thing we could improve before rolling this out more widely? (text box, probably the most important question)

TIP: MS Forms tells you how long it took on average to fill out the form….. if it’s 2 mins 13 seconds, tell the next group, you’ll see engagement jump.

Protect against the day it goes wrong

Before turning on enforcement for any wave, ask honestly: what will I do when something goes wrong?

Set up emergency access accounts, Temporary Access Pass, a plan for lost or replaced phones, and a way to help people travelling with poor connectivity. Decide in advance who can roll back a policy or provide an exception and what evidence justifies it, instead of making that call under pressure with an (irate) audience watching.

TIP: Test your break-glass accounts by actually signing in with them, on a schedule, and record that you did. An untested break-glass account is a theory, and you’ll find out which one it is on the worst possible day.

Being able to roll something back technically isn’t the same as having an agreed rollback plan. Only the plan actually helps you on the day. Define your go/no-go criteria per wave, and be willing to pause the next one. A delayed wave is a decision. A failed wave is an incident.

Every exception needs an owner, a documented reason, an expiry date, and a review process. Skip any of those and the first exception you grant under pressure becomes the template for the next fifty, until the exceptions register is doing more work than the policy it’s supposed to be an exception to.

TIP: Use SharePoint lists and an automation account to automate the review process.

Pay particular attention to the service desk. They’re a key part of your security setup even if nobody has labelled them that way. Make sure they’ve registered for MFA themselves, can tell the difference between an MFA failure, a Conditional Access block and a device compliance issue, and know how to verify someone who has lost their phone without that verification becoming the easiest way around your controls.

A strong MFA rollout with a weak help desk identity verification process doesn’t remove risk. It just moves where an attacker points.

Security teams often get the credit for a successful MFA rollout, but in reality ours depended on support staff and local IT contacts who patiently helped people through their first sign-in and answered the same question over and over. Their role carried the rollout forward….. You know who you are, and thank you!

Prove what “done” actually means

“98% of users registered” only means people completed a setup wizard. It doesn’t prove the deployment succeeded.

Better indicators: legacy authentication genuinely eliminated rather than discouraged, privileged accounts on phishing-resistant methods, an exceptions register that actually gets reviewed, break-glass accounts that have been tested rather than just created, and support ticket volume dropping in the weeks after a wave rather than only on launch day.

Measure per wave as well as overall. Ticket volume per wave is the number that tells you whether your preparation is improving or whether you’re just getting better at absorbing the same problems. If wave four generates as many calls per user as wave one, something in your process isn’t learning.

TIP: Baseline your service desk ticket volume for the four weeks before wave one. Without it, “tickets went up” is a feeling rather than a measurement.

Agree what “done” means, and who signs off on scope, enforcement dates and accepted risk, before you begin. Otherwise “success” gets redefined after the fact, which helps nobody when an auditor asks for evidence.

Final thoughts

We thought we were delivering an MFA project. What we actually did was change how hundreds of people use technology every day, and that change doesn’t end on rollout day. That’s when the real work starts: monitoring adoption, reviewing exceptions, retiring temporary authentication methods, and supporting new starters who haven’t seen any of this before. Somebody needs to own that work after go-live, and you should agree on it before you start rather than discovering it afterwards.

Setting up MFA is the easy 20%. Plan the other 80% properly, run it in waves, and the technical rollout becomes much smoother. Skip it, and no amount of Conditional Access policy design will save you from months of support tickets.

This is the first part of a four-part series on rolling out MFA the right way.

Got questions? ping me on LinkedIn.