The real threat is that we uncritically adopt language such as alignment.
Implicit in this is the idea that AI is a inscrutable matrix and going to remain that way and we'll need expert interpreters to make sense of it.
We need to insist on building tech that's explainable by design.
Alignment just means "this machine operates in ways that align with the intent of its users". It doesn't imply anything about the inscrutability of the machine in question. A gun with a misaligned scope would likewise fail to operate in accord with its user's intent, and likewise with potentially deadly consequences.
The gun comes with a manual on how to use it safely. I'm sure it has some complexities, but at the high level:
> A gun is a metal tube that uses a tiny, controlled explosion to shoot a small piece of metal (called a bullet) forward at very high speed.
If the gun doesn't work as intended, you can take it to a shop and someone can fix it so it works as designed.
All I'm saying is AI should be designed the same way. Treat AI as normal tech like any other and use similar language.
Learning ML, there was a high emphasis on the error part of things as most of the course was on minimizing errors. After ChatGPT, there is a weird anthropomorphization going on, where it's all about hallucinations, alignment and what not.
We have something that is statistical in nature so there should never been any expectation of error-free results/actions. The value has always been about discerning trends or the cost of errors being way lower than any good result.
Statistical learned indexes can exist in explainable tech such as a database.
In 2017 Google was writing papers about it. Then something changed.
I don't think it was the tech. It was a realization around the power and societal impact.
That last line needs a lot of workshopping. A guillotine with instructions on the bottom of the blade conforms to your request.
If there is guillotine in the weights of the model, it needs to be properly labeled so you can look it up by name using a database index (or a graph-vector index).
It helps both the bad guys and good guys. Like responsible disclosure in cyber security, we need to have a conversation around it.
> We need to insist on building tech that's explainable by design.
You realize this means insisting on terrible tech that humans can understand right? It essentially caps human progress at some point about 4 years ago.
If you are old and happy with the way things are this might sound like a good idea. It does not to me.
Oh you want tech that helps discover new science instead of parroting existing wisdom?
There is little evidence that the RSI we are discussing is capable of inventing the theory of relativity (or the more advanced equivalent). All we have seen is pattern matching in a much larger space than humans can, with some human provided verification tech.
I would argue that human-AI collaboration with explainable tech has a better chance. Continuous learning can be done in a way that doesn't violate IP or privacy.
You are assuming that humans are capable of understanding everything. They are not.
Human comprehension sets a ceiling on progress.
It reminds me of those schools that can only teach as quickly as the dumbest kid in the room can follow. We don't want that for our entire species.
I'm sympathetic to this argument.
All I'm saying is: if you have a choice between two systems with equal power of discovery and one is more understandable than the other, we choose the more understandable one.
Limits of Human comprehension and quest for power are two different motivations that could lead to black box systems that are marketed as semi-explainable.
We need to verify that human comprehension is actually limiting progress before allowing such things and even when we do, do it responsibly on an explainable foundation.
Agreed. I would vote for the "when there is a choice" version of your argument.
I'm not sure it is always possible to verify when human comprehension is the limit though. Most of research is done an the frontier of knowledge where we don't know what we don't know.
There would need to be a great deal of nuance in any law, and nuance in practical terms tend to just mean "loophole." Still, you're right that we should try to build explainable systems where it is possible/reasonable to do so first.
The problem is that you're not offered a choice. No one is making the "err towards explainability" choice.
Training data is treated as IP. Distillation is seen as an attack.
Open data, open training based systems such an Marin are just getting started. Explainability is not a priority there.
The ones who do discuss these ideas are confrontational about LLMs and not effective spokespeople.
> The problem is that you're not offered a choice. No one is making the "err towards explainability" choice.
I'm not sure that's entirely fair. OpenAI recently discussed this at length in a blog post after some accusations around Astra and the trade-offs. The grown-ups are definitely thinking about it, and making tough choices about the trade-offs.
It's reasonable to debate whether ENOUGH is being done here, and I doubt that even the most rabid AI advocate would argue that more couldn't be done, but everyone in the industry is very much actively thinking about it.
Check this out if you haven't read it: https://openai.com/index/an-alien-mind/
They discuss recent choices they made specifically for that reason.
> The ones who do discuss these ideas are confrontational about LLMs and not effective spokespeople.
Very much this. I'm very open to reasonable debate on the subject, but 8/10 times when I try someone who is rabidly pro/anti jumps in. It turns from a debate amongst reasonable people who reasonably disagree into some kind of political/religious battle of belief systems.
I think part of my problem is that a lot of peoples careers very much depend on them not understanding it and spreading misinformation intentionally.
Let it be capped then.
Ahh yes I remember the bad old days of 4 years ago when everyone decided human progress had enough, and we would have been stuck there forever if it hadn't been for LLMs ... we didn't know how good we had it