Sympathy? It's a Microsoft company that is being ran with a consistency of a startup in early seed rounds. Their downtime is abhorrent and unacceptable as far as enterprise goes. Their engineers look like absolute amateurs allowing for such low class work it results in their customers experiencing industry leading downtime.
This is uncharitable, rude, and pretty baseless. Unless you have a lot of direct, personal information about GitHub engineers, they’re dealing with a huge spike in traffic.
Is their uptime acceptable? No. But personal attacks aren’t necessary or constructive.
Please, read between the lines. Github actions has been a mess since 2019 at least. None of the instability is new, before this unprecedented growth (for a service that's supposed to scale horizontally) the excuse du jour was the azure migration, before that it was the high rate of shipping post acquisition.
Core parts of the product, like navigating to individual files in a code review, are broken
> Core parts of the product, like navigating to individual files in a code review, are broken
I think this is a good argument to underline "It's not _just_ the scale". Adding to this, the Github Code Review experience is kind-of broken, the way comments/threads are stacked in the PR overview has not improved, pagination isn't really a thing, and these issues are age old. Hopefully, one day, Github will mature.
The phrasing could have been better, but we've talked as a team about cancelling a $100k/yr contract nearly half a dozen times the last year, and the only thing keeping us from doing so is various "compliance" issues, and to a lesser degree the friction from physically moving. That's a tenuous moat, and if I were a GH PM I'd probably want to know that the current instability is somewhere near critical mass.
For CI there are a number of drop-in commercially-available options. You can make it a staged migration, first trivially migrating the build runners, followed by the more complicated integration test and deployment runners. The CI harness itself can follow, and finally moving to a code hosting and visualization service is last.
The final step is challenging; likely the most difficult part is changing all the code references and imports. Shadowing changes would be straightforward. Training your likely 25-50 engineers to use the new code review UX would likely not take that long.
Considering the wasted engineering velocity during Github outages, it's worthwhile to do even a partial migration. Github's action runners have in my experience, been the most fragile part of the platform. Given the ease of moving build and merge queue runners to alternates, it's a no-brainer.
[dead]
You are certainly free to attempt to migrate your workloads to your own or leased data center space in your area then buy your own on prem cloud like Dell Private Cloud and deal with all of the maintenance yourself. I don’t think it would amount to $100k/yr and you would be taking on all the risk.
For most teams selfhosting on 1U is more than enough which would cost about 1200€/year for housing + let's say 1800€/year for hardware acquisition/depreciation (probably more like 300€/year if you can work with second hand hardware) so that's about 3000€/year or 97% reduction.
Of course you may not want to manage the hardware yourself in which case you'd go for a dedicated server from a reputable provider (i.e. not microsoft/google) like OVH's RISE-XL, Hetzner's EX131, Scaleway's Core-9-L (or an equivalent solution from a smaller provider) for about 4000€/year.
Now let's say you add 8k€/year donations to your distro of choice, forge of choice, and other FLOSS projects and 8k€/year for backups. You got the whole thing running sustainably with >99.9% uptime for 20k€/year or 20% of the original bill.
It would be way less and risk would be lower (It is hard to reach one 9 no matter how bad you are at it), and self hosting is not the only option.
Self-hosted is not the only alternative...
Is it a personal attack if it is not directed to a person?
Why do I feel like GitHub, without the Microsoft tax, would handle this huge spike in traffic fine?
Can you comment on that?
not constructive, but also not personal. rude but based on observed facts.
How low would uptime have to get before you'd consider holding the engineers personally responsible for it?
Why are we holding the engineers accountable and not the CEO? The company isn’t one department. Responsibility should bubble up to the top.
They dont have a CEO ever since their last CEO said that human programming "wasn't going anywhere".
Agreed. At some point to have to stop hugging your ops team and start firing them.
No, you hire more of them, you give them what they need to do their jobs, and you actually listen to their recommendations. It's not even complicated. It just takes time and money, neither of which companies want to spend on reliability until their customers scream very loud en masse.
Agreed. In fact, keep firing ops until someone gets it done. There's no way your company will build a reputation for not supporting their teams.
Sympathy negated by their 2026 Pwnie award: "Lamest Vendor"
https://this.weekinsecurity.com/microsoft-wins-lamest-vendor...
> Their engineers look like absolute amateurs allowing for such low class work it results in their customers experiencing industry leading downtime.
I gather that you have intimate and deep knowledge on the teams and the problems they try to solve there.
If I own a restaurant and buy bread from a supplier — BreadHub.
And 99% of the bread that I get is good but 1% of the loaves, they forgot to add flour. Consistently, for years, they always have loaves missing a key ingredient that I still end up paying for.
I can be pretty sure that BreadHub have a pretty major internal issue, and should probably be questioning their competence, regardless of their “scale”, and without any knowledge of the “problems they’re solving”
I’m not saying users don’t have a right to be pissed or aren’t justified in looking at other options.
I’m saying that GH is operating at a huge scale with (probably) lots of technical debt and a forced migration to new infrastructure.
I would be (and am) highly critical of leadership. I’m not going to make strong assertions about ICs without knowing their context. I’ve worked at a company with a sterling reputation for engineering excellence where brilliant ICs were kneecapped by poor leadership.
I think a lot of us, at one point or another in our careers, have worked with potato leadership that can be short-sighted or political. It isn’t a comment on the engineers.
Please don't do it. Now somebody will come and will try to improve your bread analogy. We will be discussing bread for eons
Okay, let's use a car analogy...
I see you haven't experienced USFoods and Sysco.
I don't think this is an apt analogy. The bread makeup is still there, as far as I'm aware, no users have lost any data or are missing "key" ingredients.
An outage is more like a shipping issue with the supplier, if it's owned wholly by them.
Their engineers almost certainly are not the ones making the decision to let all the new traffic degrade their service for their existing paid customers.
If you haven't read this, you really have no idea what they are dealing with. https://cursor.com/blog/git-at-any-scale
They could throttle git operations to a level which isn't noticeable by normal 'human' usage but seriously slows down automated usage. These are basically DDOS attacks.
That seems a bit crass but the underlying sentiment stands. Microsoft has more money than God. 4T valuation. When you say you’re worth that much, no excuses. Figure it out.
So your proposal is for Microsoft to invest their free cash in helping support a bunch of AI coders maintain their pet projects? Sounds like you should be a CEO!
You don't need to be a CEO to come to the conclusion that throttling to discourage excessive usage is probably a good idea.
The move would have been to avoid purchasing and operating a subsidiary wit a question mark for a budget,
unless the firm was prepared to see it through.
Start charging more, cut down on CI, figure it out. You don't NEED to give away so much if it means your service is going down.
There isn’t mission critical work being done on GH. You start charging more and the people in charge of finance at companies paying for GH will actually start paying attention. Those finance people don’t care rn because GH is affordable and helps get stuff done. There’s a threshold that exist at every company and once it’s crossed the people who can make money decisions start asking questions
Typically when a company gets acquired, they get thrown a bunch of terrible initiatives like “change your cloud provider” that distract from their mission, executives saying “more AI!” and they definitely don’t hear anyone in management say “prioritize technical debt”.
and it has been so damn slow for a number of years now, it drives me insane.