> When it went to flip into the backup, we discovered that the backup fiber had a break

Pretty grim that a life critical system wasn't designed to report that the backup fibre was unserviceable until they attempted to switch over to it.

I wonder how long it was down? Days, weeks, months?

All these cable/internet service companies are incompetent as it gets.

I am paying, $1000, $1800 & $1900 for the same service at 3 different location (20 mins from each other).

Two locations, I also have old coax lines that are still active, but not paying for it.

When I bought two businesses, I learned that they were paying for a dedicated fiber but using coax service.

At one of the location, we had fiber, and paying for backup coax and wireless. But if you turn off fiber box, it wont fail over to either one.

I wouldn't surprise it was down for weeks and no one bothered about it.

Any possibility of line of site wireless?

[dead]

Sometimes your multiple fiber paths end up in the same bundle, severed by the same backhoe. It's always a fun day when that becomes apparent.

I actually got to see one of those once.

For a major trans-oceanic backbone provider, at least 15ish years ago they had a mile or two between Detroit and Chicago where both ends were on the same side of the interstate highway.

But it's more frequent on DAS (Distributed Antenna Systems, AKA small-cell or micro-cell) networks.

Also the challenge of when fibers are leased (if that's still a thing, based on the networks I helped design I'd say 'probably').

They really are analogous to Lamport's "Distributed System" quip; A damaged fiber owned by a company you have never heard of can wreck your day.

>Sometimes your multiple fiber paths end up in the same bundle

For mission critical stuff like airports I would like to think they're go for a more rigorous methodology than hope for the best on paths

lol there’s supposed to be …. But then greed gets in the way lol.

Usually it’s a combo of a tier 1 that decides that a resell agreement is “good enough” and “well just eat the loss” and then by the time stuff like this rolls around “we’ll get back to you” + a bunch of silly explanations that boil down to “you’re not gonna sue us though” start coming out lol

[deleted]

last time this happened in my area a local farmer was burying a cow and took out the whole towns connection

Nest door neighbor did this when he just decided on a whim to put in a new driveway and started digging with a bobcat. 10 mins in and BLAM severed the main Comcast coax serving the entire neighborhood that was running under his driveway.

Luckily we were on Century Link so weren't affected by his stupidity. lol

Years ago I was touring a POP in a small city. They were very proud to point out one fiber coming in one side of the building going South and another on the opposite side of the building going North.

Put all your backups in one basket, and then pray that nobody crushes the basket.

This is a pretty classic network operator story: go to great lengths to provide for physically diverse paths, then not notice when your provider refactors and grooms them onto the same bundle.

In NL this often happens because of waterways. All starts off with good intentions, multiple providers, totally different fibers exiting the building in different directions. And then they have to cross a canal or a river. It's a 50/50 at that point whether they converge on the same conduit across the water.

That is negligence, practical if not contractual.

Over ten years ago, my employer was spinning up a new DC across the state and had three links between it and the primary DC. Two were pretty direct, but we needed a third because at one point in the 300 mile path, the two main links went within 400 meters of each other.

So how national-security-adjacent critical systems like an airport system doesn't have a larger set of backup links, and validates that they are geographically separate up until the connections, and have realtime alerting on the connection status, is surprising to me.

That is fascinating. Out of curiosity, how far apart would the paths have had to remain in order to be ok with just two links? Like, I'm thinking a plane crash between those two links in the 400m zone could destory both links so a third link makes sense, but if it were 800m that would probably be fine? Or what magnitude of disaster are they trying to mitigate and how much separation would that require for just two links?

> validates that they are geographically separate up until the connections

This isn't relevant in this case. The problems occurred at different places. The backup fiber was, separately from the issues with the primary system, cut by a construction crew. It's not two fiber lines both cut at the same spot.

For all we know the contract says the FAA is responsible to notify Verizon if the line becomes dead. We don't know what the agreement is.

Could this be sabotage?

I drive by a VZ pedestal that has had it's service doors wide open for about 3 weeks.

The punchline is that it's about 2000 ft from the local CO, where (I believe) half the town's lines terminate.

There are 400,000 to 800,000 utility strikes a year in the US. Sure you could bury (!) an instance or two of sabotage in there with an "oops".

Recently, I tried calling 811 before digging in my yard. The webpage was broken and the hotline kept me on hold forever. I gave up. Small wonder.

https://blackhydrovac.com/underground-utility-strikes-learn-...

That seems more than merely plausible given current tensions and the location/timing connection with NYC and UN.

"Current tensions" has been a convenient boogieman for the past 100 years, if not longer.

You're going to be hard-pressed to find any point in the American empire's life when it doesn't have 'current tensions' with someone or other.

Sure. It could also be the first sign of the invasion of the Mole-Men.

I thought The Incredibles handled them already?!

I would put my money on it.

Similar events happened in Europe disguised as thieves stealing fiber optic. This does not have any sense economically, as the value in the market is zero so... either is an honest accident and is cleared in a few days, or is sabotage

>disguised as thieves stealing fiber optic. This does not have any sense economically, as the value in the market is zero

Do you honestly think that crackheads think that far in advance?

I've seen fiberoptic cables stolen from 2 (city) jobsites in the last 5 years, once by tweakers later caught trying to sell them as scrap copper and the second thief was never caught. This happened even though the spools had big signs on them saying "Fiber Optic Cable - NO COPPER".

This used to happen more often in the past, when copper thieves mistook fiber for conductor. There were cases where critical systems were shut down due to fiber optic theft, and the thieves were caught burning the sheathing to expose the... glass.

[flagged]

[flagged]

That's very uncharacteristic, especially given all those contingency requirements (backup policies, failovers, testing schedules etc.) imposed on corporations/companies in the wake of 9/11.

You actually have to ensure that someone isn't just checking boxes the tests passed and the tests actually passed.

I worked for a regional ISP that had a major outage when the redundant fiber provided by the telephone company was cut in one place triggering a full loss of connectivity. It also caused a massive 911 outage for 200,000 people as it isolated the 911 center from the city core.

Turns out the phone company didn't connect one side of the ring topology even though they certified they did. Needless to say lawsuits abounded.

And auditing that is expensive and you always get given hell from the assholes going “why am I paying for this if nothing has gone wrong”

The amount of shit I was given everytime I went through a checklist working at a fedramp certified company working with emergency alerts on the phone system made me prematurely gray.

I could feel the barely contained seething rage everytime I told the execs that the reason this release will take 3 days and not be instantaneous like your friends releases at a faang and that information made you embarrassed at your dinner party, is because you agreed to this process contractually years ago and now it’s a crime if I just sign off on it being ok without actually checking that it’s ok.

Yea, fangers are def not used to working on life critical systems. I personally have avoided it as much as possible, but not completely, for all the reasons you listed above. No need for me to accept that much responsibility for other peoples lives when I've been able to earn a living just fine not doing so.

A lot of people don't check their backups until they need to restore.

A lot of people are incompetent.

Nobody wants backups as a feature. The feature is restore.

There are plenty of people who are happy to check off "Backups" boxes, or talk about their backup strategy, or whatever - while hoping the day never comes when they actually need that tricky "Restore" feature.

Glacier, and I think a couple of Iron Mountain products, exist for "cheap backups, really expensive restores" because the restore is going to be paid for by insurance (or by a counterparty when it's explicitly used for escrow.) So there is an explicit market for this (in the backup-of-things, rather than backup-of-capabilities, space.)

A backup without a restore test isn't a backup at all

You don't always have a copy of your hardware to restore onto. And the test's entire purpose is that you're not yet sure whether your restore will truly work. So you can't just run a backup and restore on your true prod system, because you're not sure it won't wreck it. So you need extra money to have a second system onto which you try to restore. If you don't have a lot of money, you will want to actually use your disks for storage, not to put them into a second testing server. Of course I'm not talking about very professional companies with super critical data. Just simpler smaller scale places or consumers.

[deleted]

No, even professional companies with critical data balk at this.

I worked at a company worth a few billion and the leadership balked when they told engineering they wanted a near instantaneous failover system and our department informed them that would require paying for a second environment that could be rolled over to.

It is rare to find leaders who can accept the cost of redundant infrastructure that is there for emergency backup.

What puzzles me is why they can’t accept it when they are perfectly fine with insurance costs and I can’t see much of a difference between the two when looking at a spreadsheet of costs other than possibly tax differences between the type of expenditure.

Sure, that's another category and different considerations. I've worked at an academic lab with a limited budget where we set up file servers, but couldn't afford to do anything approximating 3-2-1. We did a nightly backup of a tiny part (most important) of the data onto another server in a different building, but like 90% of the data was just YOLO (well, RAID, but that's not a backup), and that's just how it is. Disks are pretty good though, they don't die often nowadays, and when people accidentally deleted their data, it was just gone. Would have been cool to have a backup of everything, but even just pulling out a nightly backup from a dense server with dozens of terabytes isn't simple and you don't want to slow down the server by constantly reading just for constructing the backups. It's a tradeoff.

Redundancy has costs and those costs can be spent elsewhere like having higher quality or bigger disks, or a faster network switch or better CPUs etc.

In theory, it would also be better to own two cars instead of one, because what if the first one gets in an accident or just breaks down. Yet, not everyone can afford that. Should you just buy two half-as-expensive cars than what you can buy one of, so you can say you have a "backup"? Likely the two half-price ones would be so much crappier that the one good car would cause you less trouble in expectation than driving a shitty one and then having another shitty spare one, both of which will constantly have issues.

I've worked on backup/failure systems since the mid-90s, and I've found there's one universal truth: If you don't fully test your backup/failure system, you don't have a backup/failure system.

There's generally two wrong responses: (1) We spent a lot of money on 'blah blah blah', a lot of other companies use it, so yeah, we've got a backup/failure system. And, (2) inadequate testing - either, we tested 1 of 50 services, and it worked, so the whole system can be restored; or, we gracefully tested, and it worked, so it will obviously work during not-graceful incidents.

And the root cause of this is generally that no one gets promoted for implementing an adequate backup/failure system, or it's extremely rare.

Besides what the others have said, speaking from a neteng perspective, sometimes backup lines (and the core infrastructure in general) are engineered in such a way that it can't easily be tested properly without taking other things down, or manually rolling a truck specifically to test it in isolation with extra equipment.

Not saying that's what is going on here, just that it's possible.

[dead]