It is a bit incoherent, but it's also... provably compliant in a scorched-earth kinda way?
What are the alternatives?
A regulator saying no? Easy, your headquarters has just moved the Cayman Islands, Ireland, or Switzerland... the US office is just a subsidiary leasing the brand IP and doing marketing.
Or, you just ignore the regulator behind closed doors because you're part of some black budget. You wouldn't be able to talk about that closet back there even if it did exist.
Or, you don't do any of these complicated loopholes and you simply move all the training and inference to a different jurisdiction.
If the only winning move is not to play, how do you ensure everyone stops playing?
Simple: you must destroy the game.
It does have a kind of mad logic to it.
(That is a bit dramatic, but the idea is if we accept that AI research will happen wherever there is capital available, instead of restricting the research you restrict the return that capital can earn. And the only provable/market way of doing that is to require the results to be open. In other words, if the benefits of getting all that cash to buy GPU's gets handed out to everyone for free, then there will be far less cash floating around to buy GPUs, slowing down the process)
Or you are China's Communist Party. You download the weights and use them as the starting point for your closed-everything AGI with Chinese characteristics while the US ensures that everyone else stops playing.
What stops the US government from doing the same thing?
Every chinese frontier lab releases their model weights and often more under FOSS terms (or more restrictive but still generally open terms).
If models for public inference use are required to be open weight in the US give or take some amount of limited fine tuning, then the US and China would be on the exact same playing field. Frontier labs would be pushing functionality for functionality's sake and would either be supported by govt funding or would be supported by domestic inference providers.
And then at the end of the day everyone is just either doing open research, is an inference, fine tuning, and training provider, or is the govt.
It's easy and straight forward and pushes everybody in the market towards common standards and interoperability.
i mean there is the whole capitalism problem to contend with in terms of getting support from congress ... but gotta admit, it has that "it's so crazy it just might work" feel to me too
Inferior quality maybe was accurate 5 years ago. Now they're taking over the global automotive market. Only thing preventing them from pulling another Japan-in-the-90s against the US auto market is americans' love of oversized trucks, tariffs, and import controls. If we had a true market economy, we'd see more BYD's than Teslas (and in most global markets that's already the case).
Yes, you can still get cheap Chineesium crap off Amazon for pennies on the competitor's dollar, but it's no longer necessarily true that Chinese = inferior.
> We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve this problem. While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models . However, our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs unforced).
Because the models are trained on hundreds of billions of user conversations, across more than a billion different humans. The conversations are anonymized and not easily traceable back to a specific user.
It's unknowable and not possible to prove if any one specific conversation contained the insights for solving Navier–Stokes.
We also don't know if the authors unintentionally provided data to OpenAI through alternate means, such as via alternate accounts or model feedback queries.
Anonymized doesn't mean there's no way to know whether it is in there. My ballot is anonymized, but it's known to be in the box because a checkmark was put next to my name when my ID was verified. OpenAI can trivially check their account settings to know what happened to their chats. The fact that they are being vague about this likely indicates that they have already done so and discovered that the data did go into the training set.
Further, given that this is all in the open now, they can search the training data. No way somebody is using some specific unique cutting edge mathematical approach to solve a fluid dynamic problem 99.9% of people have never heard of and it's not locatable. Considering they spent $15,000,000 already on this, they could afford to grep around to be able to state that their hands are clean.
No, it’s genuinely impossible to know how much of Buckmaster’s Codex data is in OpenAI’s training set.
First, the conversations are anonymized, so there's no simple way to inspect the training dataset and identify which specific conversations belong to Buckmaster.
Second, OpenAI uses these anonymized chats to generate synthetic training data, i.e. they fabricate new conversations based on specific conversation patterns where the model performs poorly, and uses these synthetic conversations as training data for future models. The synthetic data could potentially contain some of selections of Buckmaster's original chats, but it is unknowable how his specific writing could have influenced these synthetic data sets or what portion belongs to him. This information is untraceable and effectively double anonymized.
Third, OpenAI explicitly uses user feedback (the thumbs up or thumbs down ratings), as RLHF to train models. However, this feedback is anonymized and stripped of user identifiers. It's not possible to trace a specific feedback to Buckmaster, nor do we know if Buckmaster ever used this feature. I doubt Buckmaster recalls or can provide a list of every time he used this feature over the past year. OpenAI doesn't have one.
Note that the first and second only happen if Buckmaster "Improve the model for everyone" setting enabled, which I find unlikely. But that doesn't exclude option three from this list.
Knowability and likelihood are almost orthogonal here. If I commit a crime and perfectly destroy the evidence, my deed may be unknowable. That doesn’t make it more or less likely.
> not possible to prove if any one specific conversation contained the insights for solving Navier–Stokes
It may be. We haven’t seen the researchers’ transcripts. We don’t know what Buckmaster or his co-author uploaded to OpenAI or with what permissions (or if OpenAI actually respects those toggles).
I think OpenAI desperately wants to make a blanket denial that they didn't look at or train on Buckmaster and Alpöge's chat transcripts, but know they cannot, because the data is anonymized.
The fact they can't make a blanket denial triggers everyone's bullshit detectors, and they're getting eviscerated over it.
Again, see the Apple lawsuit. OpenAI has never been constrained by facts in what it can and can’t say.
To the extent anything is setting off my bullshit detector, it’s in the idea that this time is different (Moreover, the idea that we should assume this divergence without evidence.)
OpenAI doesn’t have the benefit of doubt. They shouldn’t for anyone who’s honest and reasonable. That doesn’t mean they’re automatically at fault. But when the twentieth person comes forward and says a pattern is continuing, I’m giving them the preliminary benefit of doubt. It’s a bit extreme to conclude based on that. But it’s far more baseless to swing to the other side and claim we need to clear the table for a serial offender.
I understand that AI is not just cut and paste, but some documents will have more influence than others w/ power law scaling. I would be very surprised if this distribution were not extremely steep for arcane math
Their base model must have been trained with hundreds of trillions of tokens several months ahead, at this point of time, it is impossible to rule out the possibility the model had seen that session at one point of time, and it probably did, without any OpenAI personnels actually know about it.
That is correct. It is possible they didn't opt out and given the timeline and anonymization of data unclear whether a particular conversation would have made it into the training set if they hadn't.
It might be easier, but I experimented a bit, and the prompted writing always felt a bit heavy handed; it tended to leak that mentioning the product was prompted. You could probably get it to work well, but it's trickier than it should be. For ads, I think you want something a bit like Golden Gate Claude, if anyone remembers that experiment:
Sure you can burn them, doesn't mean you're accomplishing anything from a thermodynamic point of view.
Metric to look at is EROI - Energy Return on Investment. How much energy do you put into growing these crops to get a unit of energy out?
Corn ethanol in the US is stupid because it's a 1.6:1 return at best, often a 1:1. To get a unit of energy out of corn ethanol, you need to put almost a unit in. And that's economic inputs, we know the sunshine is free. Sugarcane is a bit better, but you can only grow it in the tropics.
For reference, solar tends to be something like 4:1 - 10:1, wind is 10:1 +/- onshore or offshore, fracked oil tends to be down around 5:1-8:1, and a saudi well is (or at least used to be) 40:1.
Once you have a sustainable process, these last holdout areas like dredging (where you can't use batteries or wires) can have relatively terrible EROI as long as they're a small minority of the economy. All aviation and shipping combined is about 6% of final energy, and with major redesigns for shipping it could be as low as 4%. So you'd trade a much better EROI for 96% of the economy for a terrible one that only affects 4%.
You are looking it from the wrong angle the way I see it.
First there are significant quantities of oils grown used in cooking and discarded, in an often environmentally damaging way.
Reclaiming and burning these adds value to a waste product.
What is the EROI on batteries surely under 1.
These oils can provide energy storage and transmission in applications where electricity is ill suited, jet engines (coconut oil is especially suitable for this) and shipping come to mind.
> significant quantities of oils grown used in cooking and discarded
What is a "significant quantity"? Significant in terms of volume relative to a single restaurateur? Or significant in terms of megajoules as a percentage of say, the shipping industry? (Or we can be generous and just narrow it down to the dredging industry).
Put another way, what is the recoverable megajoules of all of the Netherlands discarded cooking oil in a year as a percent of megajoules needed for all the dredging the Netherlands does in a year? My gut tells me it trends towards negligible.
All for recycling, but biofuel refiners already specializing in this do at most 5-20% recycled biooil : traditional oil blends. Don't know why they don't sell "pure" recycled biofuels, but it's probably a mix of availability and performance needs. Should always question yourself: "if it's so simple, why haven't the people already doing this commercially figured it out yet?" You'll truly understand the space if you can produce a good answer to that (and you'll have a business model if you can enunciate why their constraints or limits are wrong).
----
> What is the EROI on batteries surely under 1.
EROI is used for energy sources. Batteries are an energy store. You can still do a similar "how much energy can I get out of this vs. embodied energy to create it" analysis, but energy out is measured over all the cycles in the useful lifetime of the device. It's a slightly different concept, and that metric is called Energy Stored on Energy Invested (ESOI). Lithium ion are about a 10:1 - 30:1.
EROI is also why we really need to get off of fossil fuels long before we actually literally run out (i.e. now would be great)
As a resource reaches depletion, there are two forces acting on EROI: 1) the increased cost of each unit of energy input, and 2) the reduced yield from that energy input.
The EROI curve bends negative faster than you think by looking at history.
I don't have a strong argument for or against using MCP, it honestly comes down to familiarity.
In my own opinion, a CLI tool is going to be much more familiar ground for engineers -- I wouldn't expect the 200+ engineers at my company to all have read and understood the paradigms of the MCP protocol but I _would_ expect all of us to have a strong understanding of CLI tools and what a good/bad tool is.
It's not disqualifying, but not having a linkedin is a signal that could drop you towards the bottom of the resume pile. Companies have problems with fake applicants, and ATS systems do create "how likely this person is a human" scores. Having an online identity is one signal, and for better or worse, that is often a LinkedIn history.
Like a fico score, you can't escape the game if you want to play.
Right level of abstraction is a good way of putting it. It's basically like creating a custom DSL, but flexibility of LLMs allow the DSL to be ad-hoc.
At what point will you need formal rigid syntax? Or is not having rigid syntax the point? If the latter, how much "informational noise" or ambiguity can you inject before the "DSL compiler" gets confused?
Scaling is another bit. Convertible Psuedocode a great pattern for writing functions, but is it useful for writing modules? If you're writing a paragraph to change behavior of a function, you're underutilizing LLMs. Paragraphs are best for spec'ing modules, and the LLMs already fill in the blanks. Not sure if it would be faster to psuedocode the entire module (although maybe just the interface would be a sweet spot...)
Yeah exactly. The module/directory level is currently untested. I'm working on a desktop version so I can talk to a file system, and then I'll be able to explore those problems.
My guess is that if you simply write `use some_fn from $repo/some/path`, the LLM _should_ be smart enough to infer in most cases. But we'll have to see how reliable that is.
I'll take the "best way to elicit a clarification response on the internet is to state the opposite confidently" bait...
The example listed in the article -- fanning out a few simple get-population, get-timezone, and make-summary calls -- is, in fact, useless overengineering. This is a basic promise chain with extra steps (priced with tokens).
But as with all software pattern learning, we learn the concepts with simple toy examples that generalize into something bigger. It's the generalization that matters here.
This is talking about a few methods and tricks for spawning effective subagents (collectively, that's the "harness"). Those tips and tricks are nice, but to not be considered useless, we need to make sure we understand why spawning subagents is useful in the first place. Yes parallelism is nice for some tasks, but that's not really what this is about.
The real reason is protecting your context. Yeah, we have 1M context windows that can fit all of LotR in it, but these machines work better when they're narrowly focused. Large context windows run into attention issues and forgetfulness ("Yes, you're right, it was stated I should/n't do X but I ignored it, my bad."). So subagents come into play when you don't want all the tokens associated with a subtask to pollute your main/primary context window and degrade task attention. Split that off to a subagent, let that context navigate the details, and just make sure your main one gets just the input/output blackbox results.
The trick is getting a sense for when the complexity of the task warrants that kind of context protection, vs when a single agent is good-enough. Your toy example will never have enough complexity to warrant the setup, but you might one day find a generalization that may.
What are the alternatives?
A regulator saying no? Easy, your headquarters has just moved the Cayman Islands, Ireland, or Switzerland... the US office is just a subsidiary leasing the brand IP and doing marketing.
Or, you just ignore the regulator behind closed doors because you're part of some black budget. You wouldn't be able to talk about that closet back there even if it did exist.
Or, you don't do any of these complicated loopholes and you simply move all the training and inference to a different jurisdiction.
If the only winning move is not to play, how do you ensure everyone stops playing?
Simple: you must destroy the game.
It does have a kind of mad logic to it.
(That is a bit dramatic, but the idea is if we accept that AI research will happen wherever there is capital available, instead of restricting the research you restrict the return that capital can earn. And the only provable/market way of doing that is to require the results to be open. In other words, if the benefits of getting all that cash to buy GPU's gets handed out to everyone for free, then there will be far less cash floating around to buy GPUs, slowing down the process)
reply