I'd say, just get know what you need to monitor in case things start to exceed the current configuration (perhaps copp) or the supported scale (maybe process health,memory). *Pedro Martins Prado* pedro.prado@gmail.com / +353 83 036 1875 (FaceTime & WhatsApp) On Tue, 28 Jul 2026 at 08:44, Saku Ytti via NANOG <nanog@lists.nanog.org> wrote:
This is absolutely the norm. People framing it as some engineering challenge that needs consideration are unnecessarily creating complexity and concern where there is none.
Tier1s run global flat level2 at this scale, and have since forever.
Outage information cannot propagate faster than light, no matter how you. bake it.
The justification should be the opposite, you should have a strong reason not to run flat IGP, if you don't have it, run flat.
You will struggle to justify any of what is proposed here, regarding SPF time, regarding convergence time, that these can be improved by adding complexity. SPF and convergence don't even matter in any modern design, because when you converge, you converge for all single failures too, that is, on link-down, you immediately forward around the problem, instead of waiting for SPF, since waiting for SPF takes a long time in any scale, even intraDC.
On Tue, 28 Jul 2026 at 06:48, Matthew Petach via NANOG <nanog@lists.nanog.org> wrote:
On Mon, Jul 27, 2026 at 6:14 PM Dan Snyder via NANOG <
nanog@lists.nanog.org>
wrote:
It really depends on what hardware you are using, how large of a fault domain you want if there is an issue in the L2 area, and how quickly you expect it to converge if there is a failure.
To add a bit to Dan's excellent reply
When I'm working on a network design, one of the questions first and foremost in my mind is "what's the blast radius for the common failures that will happen, and where does it make sense to put blast doors in?" Failures come from many sources; bugs in automation tools, humans making typos in inputs to automation tools, hardware failures at L1, L2, L3, fiber cuts, etc. The duration of the outage also comes into play; undersea, transoceanic fiber failures can have months-long repair cycles, whereas a failed optic in a datacenter aggregation router can often be swapped out in minutes.
While it can often be tempting to simply keep scaling up a flat network as you grow larger and larger, what you'll generally find is that your fragility increases as you do that. What's more is that the increase in fragility follows a power law that grows faster than the diameter of your flat network, and that there are stepwise jumps in the fragility of your network as you hit certain boundary points.
Recognize that restoration times are long for subsea links; that costs for them are much higher, so putting in N+M or 2N pathways subsea almost never pencils out for the finance team paying f or them; that trying to find sufficient diversity that you can do N+M for a small value of M while ensuring that there's no shared fate between any of N and M so that you can survive a subsea cut without having to do higher-order traffic engineering is an exponentially increasing problem. And then stare long and hard at the beautiful diagram of your single flat L2 network that circles the planet, and think to yourself "when 2/3 of my transpacific capacity goes down due to an undersea earthquake and landslide in the strait of taiwan, how am I going to tell a flat L2 network that I don't want *all* my traffic still flowing across the few links that remain?"[0]
That is, Mark Tinka and Tom Beecher are spot-on; IS-IS will handle the number of nodes you're talking about plus another order of magnitude without blinking; and with a few knobs, you can even keep your SPF timing reasonable with round-the-planet latencies. But just because it *can* doesn't mean that's a good network design. ^_^;
So, I would encourage you to think about the size and scope of your network, not just today but where you envision it being 3 years from now, 5 years from now, 10 years from now, and start thinking "how can I start putting blast doors into my network to limit the impacts of failure, and give me higher-level control points that will let me traffic engineer in ways that go beyond what I can do with a simple flat L2 IS-IS network?"
I'll admit--I don't know the first thing about your network. It may be that you have a tightly geographically constrained network, so your edge-to-edge latency is low double digits, you have plentiful access to diverse fiber everywhere you need to go so that you can overengineer your Quantity Of Service to a degree that you can leave everything to IS-IS to route around failures without ever needing to traffic engineer your traffic, and you have trucks ready to roll at a moment's notice to swap out any failed hardware in less time than it takes a sleeping network engineer to wake up, find their glasses, find their phone for 2FA to log into the network tooling to figure out what broke, why they got paged, and how to route around it. If so, that's awesome! In that case, you can rest easy knowing that a flat IS-IS L2 topology will scale up just fine for you, even at 10x your current size. :)
But...on the off-chance that's not what your network looks like, I've found that thinking about failure modes and how to limit the impact of them goes a long way towards developing a robust network design, and that the network design you arrive at is likely to not be a simple single flat L2 topology.
Thanks for asking the question! :)
Matt
[0]There's a reason many networks do continental IS-IS networks, and then use BGP plus their favorite flavor of traffic engineering overlay across transoceanic links. BGP provides very, very strong blast doors to limit the blast radius of IGP goofs, as well as an easy way to ensure your IGP won't continue to try passing all of your traffic across a small number of remaining transoceanic links during a major outage. _______________________________________________ NANOG mailing list
https://lists.nanog.org/archives/list/nanog@lists.nanog.org/message/2ESUIWF3...
-- ++ytti _______________________________________________ NANOG mailing list
https://lists.nanog.org/archives/list/nanog@lists.nanog.org/message/QFDLIZ5I...