The paper explains absolutely everything as if it was a tutorial "how to made your own modern agentic LLM". They even tell how they made their dataset. https://aleph-alpha.com/downloads/tech-report.pdf ; It's the first time I see this level of openness.

I worked on Kolibri, in particular pre-training data and mid-training. We strive to be as open as possible. Glad you like it.

How do you cleanse the data at this scale?

By various forms of deduplication (exact, fuzzy, substring), heuristic filters and distilling quality classifiers that annotate our data. Synthetic rephrases can also be considered a form of cleaning/getting more out of existing noisy data.

We have a lot of details in the tech report if you want to go deeper.

Hey! I’m curious if you tried comparing luxical to model2vec classifiers for the pretraining.

I’m one of the authors of model2vec, and working on training classifiers for this. I think model2vec could be better, but I haven’t had the opportunity to try this at scale. So if you did, knowing about it would be helpful!

Are there plans to make much larger versions of this model? With 500B-1T params for general purpose knowledge tasks, similar to the current leading proprietary models?

Stay tuned, this is just the beginning. Merging with Cohere will help in scaling, too.

How exactly would an open project do that?

How much should German authors be paid, $3,000 per book, like Anthropic paid?

Why expose yourself to this liability?

In my opinion not being open about which data is ingested and trained on, and trying to make that a repeatable thing for a third party, is not worth being called "open". Glad they did that.

to be fair, this discussion has been had numerous times here and the industry has arrived on "open-weight" to describe the practice of releasing the post-training weights in an open manner but not releasing the data it was trained on.

That's what they call this and I think it's a pretty clear definition these days to people in the industry.

How, is the training data public?

Isnt OLMO basically open like this? My understanding is people have recreated the model from the same data sources with repeatable responses, or reasonably close to the original (since LLMs never answer the same).

LLMs can be configured to basically be deterministic. It's only really bad for performance because language is not deterministic.

The paper mentioned is here: https://tej.as/blog/aleph-alpha-kolibri.

(This comment was originally posted to https://news.ycombinator.com/item?id=49943034, but we're merging the threads.)

I wish universities would take it upon themselves to curate the training sets for these models.

[deleted]

This alone makes it much more valuable than many high-profile releases despite not quite performing at the same level.

Yeah the pdf alone is awesome as a learning tool.

Such a crazy change from the times of Luminous, when they published a three pager with a claim that the model is similar good as „gpt 3“ (which??) with some graphs without y axis.

Bravo team!

Hopefully this becomes the new standard.

It’s seemed crazy to me that anyone thought these could stay closed or even SHOULD be closed source.

Be cautious what you wish for. Tools don't tell you what to do with them.

Open source LLMs "democratize" access to the "intelligence booster" that is AI. But while that has several benefits, it also has several downsides.

Humanity has the serious problem of being underdeveloped in the "spiritual" department. Ethics is often considered some sort of lifestyle choice, but it's actually the difference between order and chaos in a society.

Everybody being able to do anything means somebody will be able to do something you don't like. At an arbitrary scale.

Of course openness is only worse than leaving everything under the control of a select cabal of you believe that cabal to be more ethical than the rest of us.

The select cabal who believe in an eschaton they're actively trying to bring about, as well.

Would you apply this reasoning to the proliferation of nuclear weapons?

edit: why is the parent rationale sensible for AI and not nuclear technology?

Absurd and incomparable.

Why? That seems like a shallow dismissal. Why is a world changing technology okay in the hands of a small cabal in the one case and not another?

"Why is the wheel not okay to gatekeep but the nuclear bomb is?" Even the framing is manipulative from the beginning. By asking this question you're already assuming they're in any way comparable. They are not even remotely equivalent and the entire comparison is utterly absurd.

AI is more like a nuke than a wheel. Do you disagree? Can you suggest a less manipulative framing? I am personally a proponent of open source AI and models but I found this cabal framing strange when we do indeed rely on this kind of control for other world-altering technologies. And AI is different in that it enables technological development in ways quite unlike the wheel in a general sense.

Dude ChatGPT is not the equivalent of a device that can level a major city killing millions in a flash. It is self evident. This entire discussion is ridiculous.

Nuclear weapons are a wholly unique threat to mankind.

That's a misconception and manipulative framing on your part.

While "some chatbot" isn't the problem, general intelligence superior to humans absolutely is.

AI allows anybody to enact essentially anything. And your "level a major city" is just a small task really. The problem there is your lack of imagination, not the actual impossibility of that task.

The main thing that bugs me about the risks discussion around AI is the lack of specificity. Commenters here have a good grasp of the risks around finding vulnerabilities faster than they can be patched. That's good and it matches the applicability of LLMs to coding.

But the applicability and the ROI of LLMs for other use cases than coding is a lot squishier. Also correspondingly the risks are unspecific.

As for what to do, ethical disclosure of vulnerabilities provided a good framework for disclosing software vulnerabilities discovered with the assistance of LLMs. What is going to be novel and calls for our spiritual development in other domains?

I find it odd that more people don't realize that we are talking about risk of elevated *general intellectual-domain capabilities* as a resource. To be clear, I'm not claiming that LLM + RF is necessarily THE technology that poses the risk, the risk is in recursive self-improvement and whatever technologies will result.

It's very clear to that any specifics couldn't capture the risks, because the capabilities, including the risks, are one level higher than any specific techonolgy. It is the process of advancing technology itself, in accelerating speed, that poses the risk.

The reason I use coding and vulnerabilities as an example is that it is a concrete example. It is what people pay for now when they buy AI. And the risks are specific and can be examined in detail. Some threads on this board currently show that even these more concrete and specific risks are often overblown, with LLMs finding low risk bugs and sucking up resources to evaluate and fix them.

Here you are claiming that AI products are going to reach AGI or RSI in the foreseeable future. Of course you can't "capture the risks" with specifics because those are inherently unspecific futures. It's a bit like saying when we invent antigravity all hell will break loose.

I would believe those future risks more if there were a progression of risks. What other than finding vulns has those characteristics?

What poses the risk is the combination of abilities past a certain point enabling you to do basically anything.

While being unable to judge whether you should in the first place.

ai cant move atoms at unlimited rate and also have limited energy. so your claim that "basically anything" is a bit of a stretch.

That's what you believe, but you might be wrong.

You're right, people weirdly lack imagination on what "higher intelligence" (minus ethics) actually affords you, let's have a look:

What do average people currently want? They're taught, the most important thing was being rich. So they will ask their AI to make them rich. Most real life ways to get there are "sketchy" to say the least, usually downright unethical and anti-social, but US society turns a blind eye when the "Wolf of Wall Street" comes out on top and the schemes don't easily fit into average people's abilities of moral judgement.

-> Large parts of US society suddenly engaging in all kinds of "semi-legal/hyper-illegal" fraud schemes, at the expense of already saturated environmental and societal resilience. Guaranteed collapse.

Or, let's get rid of those pesky neighbors/wrong-colored people/annoying opinions? Again, "legal" is a pretty squishy concept and only really applies when you don't have the legal expertise to get around it. Now you can.

Or, look at the basics: what is "real"? You only "know" because you trust certain people and institutions. Generative AI can help with that /s.

It's not only about "building weapons of mass destruction". It's about doing the same shit as usual, but a thousand times faster/amplified. Look up poly-/metacrisis for starters. Going faster with AI when there's a wall in front of you isn't the best idea.

This is analogous to the problem of spam, which is a problem about five minutes younger than email. Before spam you had to buy ads in the back pages of magazines you think target vulnerable demographics. Meta already spews fraudulent ads in horrific volume.

In other words, ambitious frauds have already explored all of the angles and bought all the ads. At worst, LLMs will create a few more successful but less ingenious frauds.

You compare to laughably irrelevant things why?

"Ambitious frauds" haven't "explored all the angles".

You imply "LLMs" to be and stay less intelligent than humans, in particular yourself. You're mistaken.

Whenever classic p(doom) sentiments are explained through a text with obvious LLM markers, I wonder whether I am looking at a superhuman persuasion attempt.

..or maybe the commenter did look at superhuman persuasion long enough to believe it would be best to channel those ever the same fear fantasies from the LLM through their account to the reader.

On a more serious note, just look at the doom premises here: "Large parts of US society suddenly going criminal" is from the movie "The Purge", I think. It is fiction.

The idea that generative AI takes away our ability to find out reality. ... I don't know. People write about that a lot, but it still seems very far fetched.

Maybe through some terminally online overconsumption, like with social media? I wouldn't know.

With new AI capabilities we will have to adjust, I am sure. Media, science, education and law are changing very visibly right now. Those p(doom) narrations just seem to be pre-IPO hype though.

It is just so so dangerous. That is why they want to go public and only want to care about optimizing for the next quarter ...right before breaking into AGI. /s

Ah, didn't think of this. Some rogue militia group might try to use this LLM (or create their own LLM based on this work) to help them create biological weapons or to do a mass hacking the infrastructure of targeted country.

Wonder what safeguards Kolibri uses to prevent this? Or if they even can

I don't think that AI is as much of a boost to bioweapons as people think. Lab work doesn't get easier just because the experiment design part does.

You can simulate things.

Given enough smarts, you can simulate anything. Quantum AI is an active goal.

Chemical and biological warfare is hard. A cult in Japan created a mass casualty event using nerve gas. Which is the only somewhat "successful" terrorist WMD attack I know of. Knowledge of how to create these weapons isn't new. Guns and bombs are the most widely used terror weapons for a reason.

Enter LLMs to help through all those pesky hard parts.

The hard part is probably finding the lab equipment and chemicals without being noticed, and choosing not to use it to manufacture drugs instead which would be much more profitable

Intelligence was never the bottleneck tho

For terrorists it is tho. Four lions is basically a documentary. The only clever and creative terror attack happened 25 years ago.

Didn't Mexican cartels kidnap telco workers to build them separate infra? We expect terrorists to be less resourceful?

What's stopping them from using claude today? What could anyone possibly do to stop them from using open models? This line of thought seems like corporate/political/pr pandering more than a meaningful concern.

Oh definitely, and I’m in the cybersecurity space so I’m already on the “worst case scenario committee” hah.

But the alternative just seems… so much worse to me?

A select few groups gating access to the ability to do everything seems like neo-fuedalism in the making.

And to be fair even the gating that we do have (daybreak, CVP, etc.) is already being circumvented via keys being stolen and sold on the dark web.

Yes, a "select" (rather, self-selected) elite "controlling" AI according to their wishes, what could go wrong?

Clearly not a "better" scenario. The real problem though seems people feigning helplessness? You can't leave society "to its own". You are part of it and go where it goes. So better start steering.

When access to AI gives you abilities you cannot use responsibly, you shouldn't have access to that. Just like you shouldn't be allowed to drive a car or fly a plane or command a rocket without proper guardrails, safeguards, prerequisites, etc.

"General" intelligence isn't present in humans, why does it need to be in AI?

> When access to AI gives you abilities you cannot use responsibly, you shouldn't have access to that.

What do you mean "cannot"? As in you are granted abilities that have no responsible use?

"Cannot" as in presently cannot. That includes the case of things that have no responsible use.

You live in a curated world and rarely or never encounter such things. Precisely because your environment is curated that way.

Look at how you can't buy WMDs. They have no responsible use for you.

> When access to AI gives you abilities you cannot use responsibly

I don't think this is realistically a problem at all. It just makes certain types of research cheaper and less time-consuming. And, again, this is also a problem with the american services.

I think the unethical things are happening already with boutique firms. I admit I don't get the concern.

> Humanity has the serious problem of being underdeveloped in the "spiritual" department.

As is shown to us by filthy rich people every day.

Or did you mean the burglar in the fawellas?

[dead]

His point about regulation and innovation is great and I wish more people thought like that.

One of humanity’s biggest problems here is we don’t know how to do moderation.

We have two modes. One is a brick taped to the accelerator and damn all consequences, driven by national pride or corporate greed or egos. The other is a brick taped to the brake driven by histrionic doomers and anti-everything pessimists.

The extremes are loud and fit in a tweet. Nuance is quiet and contemplative and usually requires an essay or a book. It’s also dynamic. Nuanced positions evolve over time as new things are learned. Extremes tend to be fixed and rigid. All this, I think, gives them higher memetic fitness in the discourse.

I don’t think this is new. Look at nuclear power, a largely pre-Internet example. You had pro nukes who minimized and hand waved away any risk and anti nukes that wanted it utterly outlawed. Nobody said “hey this is a great zero carbon source of energy but we really need to think it through carefully and manage it well.” Or if they did they were drowned out by the loud screaming extremes.

(This comment was originally posted to https://news.ycombinator.com/item?id=49943034, but we've since merged the threads, so I've moved it into the subthread which is specifically about the paper being responded to.)