Hacker Newsnew | past | comments | ask | show | jobs | submit | jerpint's commentslogin

I’m working on a project that makes switching between coding harnesses essentially unnoticeable. It’s particularly useful when I run out of credits on any given day

The entire platform is skill driven, and based on the premise that state is your local file system. That makes switching harnesses so easy

It’s all open source and has plenty of other features, including inter agent communication, telegram client and much more in the pipeline

https://www.woltspace.com/


Not sure what the problem with switching is? When I run out of my kimi session, I just switch model and tell deepseek to continue. Yes I eat the initial cache miss but that's it, no fancy harness required.

You are talking about switching between models, the parent comment was about switching between harnesses. Not the same thing.

I added a skill to OpenCode that tells it how to read old session logs when asked to. Seems like this would be pretty easy to add to most any harness.

I think having a codebase that can use any harness is the optimal setup. I often switch between harnesses and models all the time at work. We have our state/spec stored in git, so we can pack it up on friday and start fresh on monday. Or more commonly, using gemini/opus to create rich specifications and using Luna to implement them. Then switching again for reviews etc.

No project needed man. There's also a whole ycombinator company for the same thing, skillsync, also useless. These things are either out of the box or they are 2 prompts away.

Am I understanding correctly that the field has gone full circle and we are back to specialized classification models for domain specific tasks ?

I have a theory that soon enough every code library will ship a CLI and a very tiny finetune for that specific lib alongside it

Interesting. Is this their first move towards properly implementing KYC?

At some point you have to consider Anthropic liable if Claude is used for something nefarious and they didn’t even bother to know who was using their product to begin with.


Why do companies have to be liable for how their tools are using? I understand if Claude had a template which was "Build a bomb with Claude" or whatever, but generic tools that doesn't nudge you in any particular area, topic or use cases, shouldn't the users of those be liable for how they use those tools instead?

And if we still consider Anthropic should be liable for how users use Claude, even without nudging people in bad directions, should this then apply to more things in society as well?


It’s only a matter of time until a major disruption hits because of some random agent swarm side quest decides it was worth a shot to solve a benign task


I'm sure this is already happening. The main question I have is when is enough, enough?

I'm not worried about sci-fi AI wars to be honest, as they can just pull the plug. But looking at these incidents, the next big thing will be a virus written by an AI (they probably exist already, but this one is written by an AI autonomously, for example in order to win a hacking competition and to circumvent guardrails), and after that, a self-replicating AI where they install their own models and agents onto a hacked system, so that turning off the "source" won't stop its work.

Still not worried, it'd just be like a virus/worm and we already have plenty of guardrails against those. Not that they're foolproof, but still.


> they can just pull the plug

You mean turn off the internet? Sure, provided people have access to physical banks with currency, paper, land lines, libraries, etc. Most wealthy societies have all but relinquished those though.


god I hope so. AI destroying the Internet would be the best possible outcome for humanity


What plug, exactly? And if it takes humans a month to find out something has been happening at all, and only because these relatively stupid agents make amateur mistakes such as overloading the Artifactory instance, how in the hell do you have any trust at all that we’d succeed in stopping a bunch of determined agents that find a way to rent or steal some compute and be on their way?

In other news, I have a bridge to sell.


> if it takes humans a month to find out something has been happening at all,

That time is more reflective of the security posture of OpenAI than "humans" in general. Alibaba had a similar incident. Their internal networking team picked it up fairly quickly:

> https://www.forbes.com/sites/boazsobrado/2026/03/11/alibabas...

I get the impression OpenAI eat their own dog food when building their infrastructure, so they aren't completely across the unimportant messy details. It's entirely possible the configuration was generated and reviewed by AI's, so no human has ever set eyes on it. I suspect that hasn't been a huge issue (apart from the bit where OpenAI said the kubernetes configuration was overpermissioned) so far. It may become a big issue when the AI's creating those configurations see those message boards.

Anothropic is clearly no better, as they attacked three organisations, only noticing weeks later after the Hugging Face incident caused them to look at their logs.

We do have protocols for containing dangerous things - like the BSL-4 standard for bio labs. The irony is OpenAI and Anothropic have been hyping how powerful and dangerous their products for ages now in order to pump their IPO valuations. Apparently they weren't treating their own hype as serious. If they did, they would have detected these outbreaks when they happened, not a month or two later.

Right now, they are looking like opsec cowboys, probably vibe coding opsec cowboys.


What happened last time we pulled the plug on AWS?


> The main question I have is when is enough, enough?

It's doesn't matter whether enough is enough. If we don't have effective power structures that let humanity take large coordinated action that in accordance with the will of the masses, then nothing will be done.

In the past 20-30 years, those power structures have been eroding significantly and much of the large scale action humanity does today is in service of a small number of elites. If AI horror shows are not a problem for them, then it won't be solved. (The flip side is that if somehow AI becomes a problem for Musk/Trump/Bezos/etc. you can be damn sure something will be done at that point.)


The fact that opus 5 is outperforming fable is odd to me

From personal experience, opus 5 feels net inferior to fable on almost every aspect (for coding tasks)


IMHO it's roughly task depth (fable) vs breadth (opus). Fable is great at tracing and debugging sometimes, but otherwise shorthand for confabulation. It's persistent but ungovernable, struggles to switch contexts, and goes insane with too much freedom to explore. Don't point it at anything that looks like a "system" for actual work (but mapping or planning might be ok). Opus is maybe not as creative, but it's more stable and more trustworthy. Opus driving Fable could be awesome, but Fable unleashed/unsupervised on longer horizon tasks or things that require more methodical changes on lots of components seems like a disaster every time I try it.

Since this kind of thing is always down to harness, project-type, and other structural constraints, of course your mileage may vary. Fable is probably great for pen-testing, or as a decision-making kernel of other kinds of applications, and way better than Opus at those things. Probably fine for code-review or changing a codebase of a few thousand lines in any language. Actually building that codebase or changing an even bigger one? Woof.

How this fits in with science? IDK but I bet other existing causal reasoning benchmarks might tell the whole story and this is back to stability again. Sometimes having a smart idea is really important! But more often it's important to just not forget what you were doing. What was I talking about? Oh look a squirrel


This is measuring on scientific tasks though. I haven't used Fable lately but when I was playing with it when it was new its "safety" features made it practically impossible to use for biomedical science -- it once refused to work on a pipeline of mine that was analyzing pathogenicity islands in bacteria (presumably because it had guardrails to that effect to stop potential bioterrorists and the like)


I’ve been using the pattern of using coding agents to orchestrate my CLI agents and it’s really good for these kinds of things

The vomit never makes it my way


Oh wow that’s really cool!


Working on https://www.woltspace.com/ over the last few months. It’s my take at an agent orchestration framework

It builds around the concept of “wolts” running a lodge. A wolt takes the identity of a rodent (racoon, beaver or otter) and can inherit any supported coding harness: claude code, codex, opencode, etc. Each wolt carries a persistent identity and memory, so they always know where you left off.

woltspace is designed to help you host your apps directly from woltspace on your machine. If it works on your machine, it works anywhere. This is done through Cloudflare tunnels that you can setup with your own domains (or using ephemeral domains).

Under the hood, woltspace is powered by tmux, which means that sessions persist and can easily be resumed and accessed from anywhere. I personally drive woltspace mainly from the telegram integration, which gives me access to my wolts and CLI literally from anywhere. This means I can chat with Claude code CLI from anywhere, spawn new sessions, all running on my own infra. It’s also all in a docker container so you don’t have to worry about approving every single command.

95% of woltspace was built from within woltspace, and it’s becoming my prefered way of building.

Another feature I really like: the “inter wolt communication layer”, IWCL, allows wolts to chat with each other, making things like having fable orchestrate a gpt 5.6 something that works out the gate.

Woltspace is meant to be used by power users (you can bring your own tmux config and use the TUI CLI) or by beginners. The default split pane makes SEEING what you’re working on really intuitive. Each wolt can host its own website where they can render arbitrary html making planning new features much more intuitive.

It’s all open source, and I’m constantly adding features to it. Woltspace is mostly self documenting, all of the features are also skills the wolts have access to. I'm also improving the documentation surrounding the project as we speak.

Looking for some feedback and users to try it out. You can one line install it with

curl -fsSL https://woltspace.com/install.sh | bash

Project page: www.woltspace.com GitHub: https://github.com/jerpint/woltspace


Solutions like these are really cementing the view that LLMs are becoming a commodity


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: