0.01 tk/s is unusable for anything, you would wait a whole day for just 1000 token of output, what is the point of projects like this?

It is fun.

Also its answering the question of what gonna happen if you wake up tomorrow and datacenters are gone. Or internets are gone.

Some people on our globe live in countries with no internet whatsoever. Of course most of them dont have Macbook with 64GB RAM either, but it's much much easier to get than internet connection or rack of GB200.

SOTA LLMs are efficiently compression of all the knowkedge humanity has built. Having ability to run it at home to extract said knowledge is important no matter the speed.

Even if somehow all the data centers are gone, it's still uselessly slow.

But also that's a pretty extreme hypothetical. Imagine the polymarket on that.

I like seeing the latest and greatest model crammed into new systems to see how it fares. To deal with the speed, one person on reddit suggested using it in an email interface rather than a chat interface.

> one person on reddit suggested using it in an email interface rather than a chat interface

Kimi Pen Pal. Bring back lettets and postcards. Do OCR, and use one of those 3D printer-like pen plotters write the model output as a letter.

Challenge would be automating the opening and OCR preparation, and the folding and mailing of the return letter. But given it's done commercially it should be possible.

Email would indeed be fitting for K3 running on a M1 Mac, as it'd take days/weeks to receive a response, which matches with my real-world emailing experience pretty well.

I had my clanker implement this idea in a standalone Rust server that speaks IMAP and SMTP and proxies your emails to an OpenAI endpoint you configure: https://tangled.org/clee.sh/posthorn

Works in mutt; other MUAs may vary.

Love the email idea

Having the right type of interface makes a huge difference.

It reminds me of when Willow Garage chose to name their bot the TurtleBot, because if they named it anything else, people would think it was fast and capable. But when they called it Turtle Bot, people just kind of liked it and were satisfied with what it did.

At the level of Kimi 3, I probably can code only about 1,000 good tokens per day, too. (thankfully coding isn't my job)

So you subscribe to the belief we won't in future find mentalism in other galaxies or solar systems which operate on mechanisms we don't understand and think v e r y s l o w w w w w w w l y ?

(note. I am not a believer in AGI)

"useful" is highly contextual. The clock of the long "now" is not useful in the sense you mean, to synchronise your wristwatch. I'm still glad it exists.

Are there any well thought through stories about what this would look like? For example, I'm thinking about like nutrient flow, decision making, energy input, gravitational force, things like that seem to govern the value and speed of intelligence.

Hard to avoid spoilers here but Vernor Vinge hits on pretty much exactly this in A Fire Upon the Deep. Though he's interested less in the hard sci-fi aspects of how/why and more on the consequences of it (story-wise).

For the opposite, "Dragon's Egg" by Robert L. Forward is a fantastic read

I think this is a slow version of the quandry behind Quantum Computing: how do you distinguish events from the noise floor? It happens in the quantum context and it would happen in the millenial timeframe completing "operations" which have to be compared to e.g. the stability of orbit around a sun.

I can only guess some post purchase remorse.

Need to justify buying an expensive rig that doesn't do what you expected.

Specifically thinking the people they could do something AI with cpu, and realizing it isn't feasible. Happened at my fortune 20 company. They had to get approvals and ofc it was useless. Plenty people tried to explain, but they were the principle engineer, and out ranked everyone.

"It's not going to work", the topic changed, and we never spoke about it again.

You can't improve what you can't measure.

Consider this like if it were the first test

'Large Language models? They can barely produce gibberish sentences, what would this tech ever be useful for?'

- bunch of people only ~4 years ago

I questioned the speed not the output quality, that is another discussion

It applies to speed too. The project paves way for more optimization at many layers overtime.

16 tokens / s is not nothing.

It's the other way around due to poor framing, it would be much easier to compare if you [the repo] said 0.02 tps.

Confusing numbers everywhere.

16tk/s... Then 3 tks per minute. Then someone else posted 0.3tk/s.

? Readme says 60-70s per token

[dead]