He didn't "flee" to Russia. His passport was renounced/revoked/disabled while he was on his way to the original destination. He got stuck in the Russian airport for a freaking long time in a small room while the govt made sure he can't go anywhere else. Good Lord, fled to Russia it seems.
I will never touch this garbage and would rather use the disk space I purchased to be used for my purposes, not the industry's pointless fad bubble endeavors.
Bookmarks died because most humans are allergic to having to do bookkeeping and cleanup tasks that are incidental to what they really want to do.
We see that over and over again, in a bunch of non-tech fields. My kids never want to clean up their toys; they will pull out new ones, but it is a struggle to get them into the habit of putting the old ones away first. My parents used to keep a notebook with every gas fill-up they made; nobody born after about 1990 does that anymore. We were taught how to balance a checkbook in elementary school; basically nobody does that anymore, we put everything on autopay and if you are diligent you check a statement or import it into Quicken once a month (most people don't even do that, they have no idea what they are spending and predictably, usually no money left over). GMail succeeded because instead of putting your mail in folders, you just leave it in one lump with Google and rely on full-text search. A lot of Zoomer computer users don't even know what files and folders are, they just use apps, which take them straight to what they want to do and don't offer things like possession of your own data.
Back in the (~early 2010s) days when Google still allowed internal innovation, there were recurrent demos produced by engineers of full-text search over your web history, though of course it was your web history as stored by Google and none of this was local. It never became a product, largely because users are too lazy to go to a separate search product just for your history, or because they're too lazy to check a separate box saying "Search my history". Instead I think web history became a ranking input to general search and it would mix in results that you frequently visited to the general results, which honestly I think was a more useful approach.
I cut out Social Media earlier this year, and it's been one of the best decisions I've made. Part of why I started was remembering that we didn't have doom-scrolling apps before smartphones. The change has been tremendous since I've been trying to go back to more "intentional" media. When I find myself distracted and wanting to pick up my phone, I realize there isn't really anything on it for me to do and am forced to just put it away. So I've started getting into the habit of keeping a book on me wherever I go. I've read 4 books this year!
There were sellers buying the smallest RAM units, changing the RAM chips, and then reselling them as the higher RAM SKUs. The sellers who do this don’t always care to use good RAM chips and may even use QA reject parts. When the unit doesn’t work correctly, the anger and RMA requests are directed back at the Raspberry Pi foundation.
For us dell has become that. Rep changes all the time, what appears to be random pricing and long unclear delivery dates. Instead I can for example go to Thomas Kren and configure my serves online for a fraction of the cost. They have what they list in stock, delivered a few weeks later. Everything is clear and hasle free.
At this point I have the feeling dell just charges you what they think they can get out of you. You push back and the price drops, I don't want to negotiate.
> Communication. Opus 5.5 communicates more naturally than prior models. Early testers found its writing clearer and easier to follow, which addresses some of the common feedback we heard about Opus 5. It puts the most important information up front, and its style makes it a better work partner over long sessions. As one early tester put it, “it writes the way I do.” In our own use, this has made Opus 5.5’s work easier to follow and check—which is a safety benefit as well as a practical one.
I think this is what I'm most interested in. I mostly moved to Astra because I just can't work all day with the Claude Opus 5/Fable writing style. I don't think Astra is a better model, but it's the first OpenAI one that seemed good enough to me. Definitely keen to try Opus 5.5 and see if this claim is real.
I really don't think it's productive for internet forums to constantly be criticizing language choice when the meaning is clear. Better to respond to the substance of the issue than word choice.
Edit: In response to the initial replies. To me it clearly means "releasing frontier models at any pace less than as fast as possible". It implies relative restraint compared to the previous state and without stating the degree of restraint.
I found Grim Fandango in a pile of discounted games as a tween in the 2000s. There was a cool skeleton in a suit on the front cover, so I bought it based on that.
Even though I didn't understand most references and sarcasm, few games have grabbed my attention as Grim Fandango did. The art work, the music, the writing. The whole game oozes of style that I've never seen replicated.
Even some 20+ years later I can almost recite most of Act I from heart. If you haven't played this game and like adventure games, this is one of the best.
The meaning isn't clear at all. So open for interpretation that it is meaningless. That's the whole fucking point. For all I know they are "pacing the frontier", or not. The fact that there's no meaning to it let's you know that it was a pointless waste of tokens and attention.
I've been working with agents all year, but 5.6 Sol was some sort of sweet spot for me. Something about how it communicated verbally and its engineering instincts just clicked for me, and I was able to somehow predict it and jam with it. Like a colleague you click with. It's the first model I've gotten attached to. I'm concerned that whatever model supercedes it, while technically better, just won't feel quite as natural to work with. And this makes me feel very professionally vulnerable to the labs. I miss the days when my crucial tooling came from companies as reliable and predictable as, say, Jetbrains.
I... don't know how I feel about this. I mean, the concept, sure, good idea. Is it reliable? Perhaps. But... I get a "buy merch" popup the moment I open the page, and there seems to be a pro option for what seems to be a very much LLM-aided app? ...IDK, rubs off a bit wrong IMO. Just my personal opinion.
The only thing the West is achieving here is that sooner or later China will be able to compete on their own terms rather than ours. It may buy some time but the end result is very predictable.
(ass. sports.txt politics.txt and business.txt are text docs pertaining from the sports, politics and business domains, respectively, and have equal size)
The test file belongs to the topic with the smallest size *.gz file.
Witten's group at Waikato uni were perhaps the first to work on this.
Also check out the Hutter prize if you are interested in this.
There's a pattern to the kinds of companies PE buys and I think it points to the real problem.
They like companies with some kind of moat that makes it hard to unseat them. Basically, companies where there is no alternative for the consumer. That way, they can inflict abuse but know there will be nowhere to run.
There are two different ways to achieve this. Monopoly and regulation. Hospitals have both government granted locational monopoly and tons of regulations that make it impossible to compete.
Private equity is the symptom, not the disease.
Until we get at the disease, new monsters will be born with different name filling the same ecological niche. It's economic natural selection played out in the environment we created.
PS5 emulation has gone from barely working to running Dark Souls at 10+ FPS with virtually no graphical glitches… in 6 weeks. If you’re not familiar with normal emulator development time, this is… quite extraordinary. There have been insane progress on decompilations and many other things in the emulator space.
Everyone’s trying to figure out how to convert this speed to product features at scale, but enterprises are like container ships. Lots of might but slow to turn. The littler companies can actually take advantage of this and produce higher quality products at much faster speed. I think you’re expecting too much in the short term and too little in the long term. AI-native companies are gonna eat everyone’s lunch, once they figure out how to actually do it reliably.
My eternal advice to anyone doing heavily LLM assisted projects is to do them a bit less LLM assisted. You should still use LLMs, in my opinion, they're wonderful tools. But I know reading this page that this text is mostly LLM-generated, and that leaves there to be little hope that much else of the project isn't.
It's one thing if your code is really truly "co-authored by" Claude. It's another thing if it isn't even really co-authored by you.
Evidently the majority of people see no problem with LLM prose, but me and many others think it is some of the most annoying crap possible. Please consider speaking in your own voice.
I might turn this into a blogpost if folks are interested, but my god there is so much clever info in that dashboard.
Here is one really neat bit:
A cutting edge training idea (for agents, it's been used elsewhere for ages) is on-policy RL, basically, it's not enough to say "here is an end to end agentic sequence (including tool calls etc.) that is perfect" you want to say "here is a sequence you might actually have generated that turns out to be correct".
Basically, it's more training efficient to improve models with small tweaks to do more of the right thing they are already doing sometimes than from some perfect oracular "this is the way" answer.
(if you've ever tried to teach humans new skills, you’ve probably noticed this too!)
When you do that, you care about how far the model you are updating (improving) has deviated from the one being used to generate rollouts (agentic rollouts for hard problems can take hours with lots of tool calls, so you can't keep redeploying every slight improvement).
Lo and behold, the dashboard literally has:
partial/avg_staleness (likely the measure of how many micro iterations the "generate answers" model is behind the "improving based on the occasional right answer" model)
train_infer_diff/new_infer/kl (a more direct KL divergence based way of measuring how differently the two models generate tokens)
How cool is that?!
And don't get me started on the clever ideas hiding behind dynsam/avg@n ...
>A man stranded in the bush in northern Saskatchewan was rescued last week after chopping down four power poles — knocking out electricity to surrounding communities. [...]
>But he had an axe and he knew SaskPower would have to check the downed line, so he went to work.
I would never have figured out the toggle for writing tools is in `Settings -> Screen Time -> Content & Privacy Restrictions -> Siri -> Writing Assistance`
I have to be honest. While this is obviously a smart and useful idea, it misses one of the core features of Jev: its confidence scores. Partial confidence could easily be mapped to fractional spaces, using unicode characters like U+2009: THIN SPACE. As it stands, this package is not harnessing the full power of Jev.
No, OpenAI did not solve the "wrong" Navier-Stokes problem. OpenAI did not solve the hardest version of the problem (unforced blow-up), but did give a solution to the Clay Millennium Prize Problem as written and understood, choosing the explicitly allowed forced option.
SciAm writes "in a sense, the LLM found and exploited a loophole in the framing of the question". This is pure sensationalism. Choosing option (C) (out of an explicit list of four options) is neither a "loophole" nor something "found by the LLM"; everyone involved knew this was the option they were pursuing.
With the grumbling out the way, there is some actual scientific content to the article: there's a strong argument that OpenAI's method will not extend to the unforced case, leaving our understanding of NS incomplete. This negative result is itself new and interesting (and predicated entirely on the solution found by OpenAI)!
> "The prediction was that software engineers would lose their jobs first, and the recent trend went the other way: demand for software engineers is higher than ever"
Hard to take anything the author says seriously making nonsensical claims like this. The software engineering job market has been getting worse every year since 2022 by virtually every metric. This is especially true at the entry and mid level. For example, computer engineering and computer science majors now have the #2 and #4 highest unemployment rates amongst recent graduates [1]
Students are smart to be cautious about the future, and it's annoying that adults with no skin in the game so flippantly dismiss these concerns without any data to back it up.
So Uber ToS requires you to accept arbitration, then, when they are found responsible for damages, they still don’t want to pay. Seems pretty shitty for the consumer.