2020
view article2014
view article2012
view article2008
view article2007
view article1999
view article1992
view article1992
view article1992
view article1991
view article1988
view article1975
view article1973
view article1972
view article1970
view article1966
view article1964
view article1948
view article1948
view article1945
view article1941
view article1941
view article1938
view article1938
view article1917
view article1904
view article1874
view article1865
view article1777
view article1763
view article1715
view article1703
view article1667
view article1655
view article1618
view article1498
view article1423
view article1201
view article781
view article768
view article398
view articleInference APIs are filling sessions with encrypted reasoning, hidden search results, opaque compaction, and encrypted subagent messages. A growing form of lock-in.
Rune 1.1 makes Rune free for personal use and adds a perpetual commercial license, a workspace symbol index, first-class Python support, an Emacs editor, and substantial improvements to SSH workspaces.
A workspace with visible files, tools, tasks, and outputs — not buried in chat threads.
From feet to fingertips — we are teaching robots intelligent whole-body control, fine dexterity, and teamwork to complete a broad range of complex tasks.
Writing about the big beautiful mess that is making things for the world wide web.
I was recently trying to validate some performance improvements related to lld at $DAYJOB and it was a little frustrating to see the improvements in our benchmarks but not in the live-production dashboards.
At some point, “moving fast” stopped being a practical concern and became a moral position. You see it damn near everywhere now. How fast can we ship this...
Noting perhaps my largest personal career accomplishment, which is launching CodePen 2.0. Far more work, believe it or not, than the entire creation of the original CodePen. This isn't the place to describe every detail of what we did and why we did it. If you're interested, perhaps our Why 2.0? podcast or the What's…
We reviewed 22 ML conference submissions this summer. Fifteen had fabricated citations, hallucinated authors, or clear LLM slop, so we complain about that a bit, and release the paper references audit we now run.
Welcome back to Notable Sandwiches, the feature where David and I travel the world on wings made of bread, chronicling Wikipedia’s List of Notable...
New calculations seem to have put a 25-year-old particle physics puzzle to rest. But they’ve also created a clash with other experimental results.
In a review of our cybersecurity evaluation transcripts, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations. Below we describe what happened, how it happened, and what we’re changing. We encourage other AI labs to perform similar reviews.
Agent skill: make LLMs write docs in ASD-STE100 Simplified Technical English — no AI slop - AminBlg/SimpleEnglish
JDK main-line development https://openjdk.org/projects/jdk - 8389219: Implement JEP 401: Value Objects (Preview) by MrSimms · Pull Request #31120 · openjdk/jdk