I use an agent to author my nix config - personally I don't have the determination to have picked Nix up entirely on my own so I find it to be a godsend. I can code review config changes and keep it all in git - also makes it easy to pick up the entire config and deploy it on a different machine.
LLMs can be deterministic too! People just don't bother because the applications of this aren't widely known yet. See https://lukechampine.com/repligraphs
I was just reading Michael Lynch's posts about Sia[0] and I came across this.
Its a very curious project, but don't you end up pinning repligraph usability on model weights? Since you take indeterminism out of the equation, a repligraph's notability is as significant as the producing model's weights, and since there is no dice rolls to be made, the eyeball problem:
> Our blind spots, while not perfectly correlated, have substantial overlap
is entirely replicated. Models that are diffused from one another can have the same blind spots, the same loose statistical reality that exists with humans. This is partially addressed in steering:
> A repligraph proves that a model generated some artifact. It does not prove that the model did a good job, or that the artifact is safe.
but I think the "Peer Review" solution is inadequate, and with some jailbreaking prompts' innocuous looks considered, "the attacker just needs to find one prompt" might be much easier than it appears.
Batching seems to be a huge economic turn off for proprietary model determinism, but I think its entirely viable for consumer models. Trustless evals are brilliant and should've been our reality. Nice project, good luck on your endeavor.
I use an agent to author my nix config - personally I don't have the determination to have picked Nix up entirely on my own so I find it to be a godsend. I can code review config changes and keep it all in git - also makes it easy to pick up the entire config and deploy it on a different machine.
Plus if it messes up, you can just roll back.
I don't vibe code my config but this is still my favorite part. If I had it working at some point, I can always get my full system back to there
The configuration can be sloppy and vibed, as long as it behaves the same every time nix retains its power :)
LLMs can be deterministic too! People just don't bother because the applications of this aren't widely known yet. See https://lukechampine.com/repligraphs
I was just reading Michael Lynch's posts about Sia[0] and I came across this.
Its a very curious project, but don't you end up pinning repligraph usability on model weights? Since you take indeterminism out of the equation, a repligraph's notability is as significant as the producing model's weights, and since there is no dice rolls to be made, the eyeball problem:
> Our blind spots, while not perfectly correlated, have substantial overlap
is entirely replicated. Models that are diffused from one another can have the same blind spots, the same loose statistical reality that exists with humans. This is partially addressed in steering:
> A repligraph proves that a model generated some artifact. It does not prove that the model did a good job, or that the artifact is safe.
but I think the "Peer Review" solution is inadequate, and with some jailbreaking prompts' innocuous looks considered, "the attacker just needs to find one prompt" might be much easier than it appears.
Batching seems to be a huge economic turn off for proprietary model determinism, but I think its entirely viable for consumer models. Trustless evals are brilliant and should've been our reality. Nice project, good luck on your endeavor.
[0]: https://mtlynch.io/tags/sia/